Anthropic makes Fable 5 safety mechanisms transparent

The company acknowledged a misstep in deploying the Fable 5 model. Previously, safety filters operated silently, leading to unexpected rejections. Now, when a safety mechanism triggers, the system will switch to Opus 4.8 and notify the user of the reason for the refusal.
In API requests, details on why content was blocked will be available soon. Anthropic notes that the silent filters were initially chosen to accelerate rollout but resulted in excessive false positives. The company acknowledges that making the safety mechanism transparent may make it vulnerable to evasion, potentially increasing false blocks as it refines its classifiers.
To address errors, Anthropic asks users to submit feedback through the feedback function in Claude Code or an appeals form for API users.
Related stories
- How model auto-switching works in Claude Fable 5
- Anthropic to temporarily limit access to Claude Fable 5 for subscribers
- Anthropic releases Claude Fable 5 and Mythos 5 models
- 1Password fills Claude's logins without showing passwords
- Claude's file checker reads C2PA credentials
- Claude Cowork gets its own browser, no Chrome needed
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
