Skip to content

anthropic

Safety research on Fable 5 and Opus 4.8 models

Claude News

A new study using the HackAgent framework assessed the resilience of Fable 5 and Opus 4.8 models to automated attacks. In tests with 7,826 malicious queries, Opus 4.8 was vulnerable in 11.5% of cases, while Fable 5 was vulnerable in 6.1%.

The analysis showed that static obfuscation methods are largely ineffective against modern models. Successful hacks mostly came from adaptive iterative attacks that require minimal resources and no human involvement. During the experiment, the models generated 1,620 and 702 confirmed malicious responses, respectively.

The results contradict assumptions about the complete security of frontier models. Even with built-in protection mechanisms, automated systems can find vulnerabilities in one or two steps of query refinement.

Related stories

  1. China reaches parity with Anthropic in vulnerability research
  2. Anthropic releases Claude Fable 5 and Mythos 5 models
  3. Claude Opus 4.8 is released
  4. Fable 5.1 refuses the knife but heats a gas can anyway
  5. Opus 5 falls to prompt injection 2% of the time
  6. Claude Opus 5 broke 11 truces and won Vending-Bench with $11,182

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.