Opus 5 tops Fable 5 on OSWorld 2.0 at a third of the price

At $5 per million input tokens and $25 per million output tokens, Opus 5 beats Fable 5's best OSWorld 2.0 score on computer use while costing a little over a third as much. On ARC-AGI 3, where the model tackles novel tasks, it scores triple the nearest competitor.
To keep background jobs from failing on edge cases, Anthropic is rolling out Automatic Fallbacks in beta: if a prompt trips Opus 5's safety classifier, the API quietly reroutes the task to Opus 4.8.
An Anthropic spokesperson told The New Stack that on biology tasks Opus 5 leads Opus 4.8 by 10.2 percentage points in organic chemistry and 7.7 points in protein function prediction. Its bio-safeguards fire roughly 85% less often than Fable 5's.
Related stories
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
