New Claude models consume more tokens but cost less per solved task

A benchmark of 10 Terminal-Bench tasks run through Claude Code across Sonnet 4.6, Opus 4.7, and Opus 4.8 shows a shift in efficiency. Each run was instrumented using OpenTelemetry and SigNoz.
In terms of accuracy, Sonnet solved 5 tasks, Opus 4.7 solved 7, and Opus 4.8 solved 8. While Sonnet was the cheapest in raw cost at $6.30, Opus 4.8 proved cheaper than Opus 4.7 at $8.06 versus $8.51.
The critical metric is cost per solved task: $1.26 for Sonnet, $1.22 for Opus 4.7, and $1.01 for Opus 4.8. Opus 4.8 consumed the most tokens at 6.06 million compared to 2.88 million for Sonnet, but nearly all were cache reads, the least expensive category.
Tool error rates dropped from 1.72% for Sonnet to 0.85% for Opus 4.8. Sonnet failures were primarily timeouts, while the single Opus 4.8 failure was a fast, incorrect answer on a complex task.
Related stories
- Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
- Fable 5.1 tops Artificial Analysis at 66, costs 20% more
- Claude Sonnet 5 increases performance at a higher cost per task
- Anthropic will bill again for requests its safeguards block
- Claude hands paid users a spare limit reset until Oct 22
- Opus 5.5 leads the index and burns 260M tokens doing it
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
