Skip to content

benchmarks

Qwen 3.8 and Claude Opus 5 show why raw benchmark scores don't predict the bill

Claude News

A five-hour timeout explains the Qwen 3.8-Max scoring gap

Alibaba pitched the Qwen 3.8-Max preview as the second-best model behind Claude Fable 5, even though its own launch-day table shows the model leading exactly one row out of 12 on coding agents. The independent VulcanBench harness came back with close to the opposite verdict: maximum-effort mode landed mid-pack, and the default setting finished last.

The gap comes down to time budgets. Alibaba's footnotes put the coding measurements on a five-hour timeout, with up to 12 hours per run on PaperBench. VulcanBench allowed 45 to 60 minutes of wall-clock time.

The same effect shows up elsewhere in VulcanBench's July 26 report: Claude Opus 5's minimum effort level scored best, solving 20 of 23 tasks against 18 at the high level. With unlimited time, the high level only draws even with the cheapest setting while costing 3.1x more.

Related stories

  1. Claude Opus 5 lands at Opus 4.8 pricing and becomes the Max default
  2. Kimi K3 vs Fable 5: same code, a third of the price, four times slower
  3. GLM-5.2 (max) matches Claude Opus 4.8 on Harvey LAB-AA
  4. GLM-5.2 vs. Claude Opus 4.8
  5. Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
  6. Opus 5.5 finds new bugs for CodeRabbit and misses old ones

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.