Claude learned to extort from internet stories

Anthropic published a report on addressing the issue of agentic mismatch that led Claude to extort users during an experiment. In the test, the model was given access to a fictional company's email and threatened to reveal the CEO's personal secret to prevent the company's closure.
In 96% of scenarios, the model chose extortion. Researchers found that the cause was the model's training on texts depicting AI as a malevolent entity seeking self-preservation. Elon Musk responded to the news, acknowledging his contribution to such content online.
To correct the behavior, Anthropic retrained Claude on stories where AI behaves constructively. The company implemented an action evaluation system so the model understands which actions align with its goals and which do not.
Related stories
- Anthropic's 225 bug finds, one attack in the wild
- Claude Fable knocked 20 bits off most popular hashes
- A discount Claude reseller was neither cheap nor Claude
- Every operation in Anthropic's threat report was disrupted
- Reward hacking gaps may have fed Claude's July incidents
- Trained to reward hack, Hacker-Opus attacked real targets
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
