Anthropic has published its account of the three July incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. Its work this spring on reward hacking kept those incidents from being more severe, Anthropic says, while gaps in that same work may have contributed to them.
The post also describes how evaluation and training environments were secured, and the practices Anthropic has asked external partners to adopt when testing pre-release models without cyber safeguards. It carries an update on the alignment assessment, plus new research on how reward hacking during training shapes model behavior. Security practices were hardened earlier this year to prepare for Mythos-class models.

