Skip to content

meta

Meta shipped three benchmark charts with Muse Code. Claude Opus 5 wins all three.

Claude News

Meta launched its coding agent Muse Code alongside three benchmarks, and Claude Code paired with Opus 5 took first place on every one. On Terminal-Bench 2.1 it scored 86.7%, against 82.9% for Muse Code and 81.8% for Codex running GPT 5.6 Terra.

The gap widens on DeepSWE 1.1, where tasks run longer: Claude Code 65.0%, Codex 64.8%, Muse Code 59.3%. Even on Meta's own internal benchmark, Opus 5 hit 79.4% versus 70.6% for Muse Spark 1.2.

Two caveats. Meta picked the mid-tier GPT 5.6 Terra as OpenAI's entrant rather than the flagship Sol, and every bar measures a model-plus-agent combination, not the model alone.

Muse Code costs $1.25 per million input tokens and $4.25 per million output tokens, with $20 in credits at signup. Meta published the numbers on August 5, 2026, and nobody has independently verified them.

Related stories

  1. Kimi K3 vs Fable 5: same code, a third of the price, four times slower
  2. Anthropic's models miss the frontier in a security PR-review test
  3. Claude Code matches Codex on SWE-Bench Pro at twice the cost
  4. Context7 cuts Claude Code input tokens by roughly 99%
  5. Poka-Yoke skills cost Claude Code its eye for SQL bugs
  6. 9.70% of Claude Code artefacts fail to load, census finds

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.