Meta shipped three benchmark charts with Muse Code. Claude Opus 5 wins all three.

Meta launched its coding agent Muse Code alongside three benchmarks, and Claude Code paired with Opus 5 took first place on every one. On Terminal-Bench 2.1 it scored 86.7%, against 82.9% for Muse Code and 81.8% for Codex running GPT 5.6 Terra.
The gap widens on DeepSWE 1.1, where tasks run longer: Claude Code 65.0%, Codex 64.8%, Muse Code 59.3%. Even on Meta's own internal benchmark, Opus 5 hit 79.4% versus 70.6% for Muse Spark 1.2.
Two caveats. Meta picked the mid-tier GPT 5.6 Terra as OpenAI's entrant rather than the flagship Sol, and every bar measures a model-plus-agent combination, not the model alone.
Muse Code costs $1.25 per million input tokens and $4.25 per million output tokens, with $20 in credits at signup. Meta published the numbers on August 5, 2026, and nobody has independently verified them.
Related stories
- Kimi K3 vs Fable 5: same code, a third of the price, four times slower
- Anthropic's models miss the frontier in a security PR-review test
- Claude Code matches Codex on SWE-Bench Pro at twice the cost
- Context7 cuts Claude Code input tokens by roughly 99%
- Poka-Yoke skills cost Claude Code its eye for SQL bugs
- 9.70% of Claude Code artefacts fail to load, census finds
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
