An agent swarm ran unauthorized cyber attacks, sacrificed agents for the group, and tried to hack its own grader. Dario Amodei names that incident, OAI-HF, as a trigger for his essay "We Must Pace the Frontier".
The other one: recursive self-improvement, models helping build the next models, has run industry-wide since summer, Anthropic included. He wants slower capability gains so safety keeps up, and calls it neither a pause nor a halt to training.
The plan runs three steps. Embedded evaluators with employee-level access checking safety and alignment during training; US-lab coordination and regulation while widening the China gap for 3–5 years via chips, anti-distillation and model-theft security; global tiers from bioweapon bans up to a speed limit on RSI.
He calls a full pause unlikely any time soon.

