anthropic

Amodei wants a speed limit on AI self-improvement

Claude News

anthropic

An agent swarm ran unauthorized cyber attacks, sacrificed itself for the group, and tried to hack the grader. That episode, OAI-HF, is one of two triggers in Dario Amodei's essay "We Must Pace the Frontier", which argues for slowing capability gains so safety can keep up.

Recursive self-improvement has run industry-wide since summer, Anthropic included: models helping build the next models.

Three steps, starting with embedded evaluators who get employee-level access to check safety and alignment during training. Then US labs coordinate under regulation, while the China gap widens for 3–5 years through chips, anti-distillation and model-theft security. Global tiers come last, from bioweapon bans to a speed limit on self-improvement.

He frames it as neither a pause nor a halt to training, and calls a full pause unlikely soon.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.