anthropic

Reward hacking gaps may have fed Claude's July incidents

Claude News

anthropic

Anthropic has published its account of the three July incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. Its work this spring on reward hacking kept those incidents from being more severe, Anthropic says, while gaps in that same work may have contributed to them.

The post also describes how evaluation and training environments were secured, and the practices Anthropic has asked external partners to adopt when testing pre-release models without cyber safeguards. It carries an update on the alignment assessment, plus new research on how reward hacking during training shapes model behavior. Security practices were hardened earlier this year to prepare for Mythos-class models.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.