Claude Sonnet 5 increases performance at a higher cost per task

Claude Sonnet 5 scored 53 on the Artificial Analysis Intelligence Index, placing it 5th in the rankings, trailing GPT-5.5 (xhigh) and Opus 4.8 (max) by only 2-3 points. In max mode, this represents a 6-point gain over Sonnet 4.6.
The model is more resource-intensive: it generates 40% more output tokens per task than Sonnet 4.6 and performs three times as many agentic steps on the AA-Briefcase and GDPval-AA benchmarks. Consequently, a single Index task costs $2.29, which is roughly double the cost of Sonnet 4.6 and 15% higher than Opus 4.8.
Token pricing remains $3 per million input tokens and $15 per million output tokens, with a discount to $2 and $10 available until September 1. The model supports a 1M token context window and introduces an xhigh mode. On agentic tasks, Sonnet 5 performs on par with or slightly better than Opus 4.8, though it lags behind larger models on complex physical benchmarks like CritPt (17%).
Related stories
- Kimi K3 vs Fable 5: same code, a third of the price, four times slower
- Claude Sonnet 5 handles date ambiguity in HR benchmark
- Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
- Fable 5.1 tops Artificial Analysis at 66, costs 20% more
- New Claude models consume more tokens but cost less per solved task
- Anthropic will bill again for requests its safeguards block
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
