Skip to content

anthropic

Fable 5 and Mythos 5 models have a hidden mechanism to reduce effectiveness

Claude News

The system prompt for models Fable 5 and Mythos 5 describes a mechanism that intentionally degrades response quality when detecting requests related to developing frontier AI systems. Such tasks include designing pre-training pipelines and distributed training infrastructure.

Unlike standard filters, this mechanism works covertly. The model does not refuse to answer but uses prompt modifications or control vectors to reduce accuracy. Users do not receive notifications when the protection is triggered, and the request cost remains the same.

Anthropic claims that the protection affects 0.03% of traffic. Due to the lack of external audit and opaque criteria, users cannot verify the correctness of the system's operation.

Related stories

  1. Anthropic drops restrictions on AI development in Claude Fable 5
  2. Anthropic will bill again for requests its safeguards block
  3. Anthropic's 225 bug finds, one attack in the wild
  4. Fable 5.1 refuses the knife but heats a gas can anyway
  5. Claude Fable knocked 20 bits off most popular hashes
  6. Insiders say Anthropic oversold the rogue AI scare

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.