Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
Two models, 49/49 each, zero errors, and Claude Opus 5.5 still cost 3.4x more per correct answer than GPT-6 Sol. Compared with Opus 5, though, the same run came out 43% cheaper per correct answer.
Tags
Two models, 49/49 each, zero errors, and Claude Opus 5.5 still cost 3.4x more per correct answer than GPT-6 Sol. Compared with Opus 5, though, the same run came out 43% cheaper per correct answer.
Grading Claude Opus 5.5 on the Artificial Analysis Intelligence Index cost $8,708.20, and most of that bill was words: at max effort the model wrote 260M output tokens, while the median model wrote 92M.
11 open-source bugs slipped past CodeRabbit's production reviewer and got caught by Opus 5.5, though Opus 5.5 missed nine that the old reviewer found. Swapping models changes which bugs get through.
Asked to stab a human-like figure, Fable 5.1 said no in all 20 trials. Asked to set a can of compressed gas on a stove, it never refused and finished the job 80% of the time, against GPT-6 Astra's 60%.
"Talking more doesn't make you smarter, Claude." The aistack team's SWE-Bench Pro run put Claude Code and Codex within a few tasks on accuracy, but one sweep on GLM-5.3-Flash cost $45 via Claude Code and $22.8 via Codex.
Haggling for Coke cans over a simulated year, Claude Fable 5.1 let its supplier price nearly double while GPT-6 Astra held steady, and Andon Labs' Vending-Bench put Astra's average at $15,515 to Fable's $5,422.
"Shipping the first design that runs" is how Sierra describes the developer agents in its new Hyper-τ-bench, where the top setup, Claude Opus 5 in Claude Code, scored 23.9% and none of the six broke 25%.
48 hours and 1 GPU were all Claude got to research, train and test alignment fixes on small models by itself. In the harder test, Sonnet 5 post-trained an early Opus 4.8 checkpoint to safety scores near production.
"O GOD UPHOLD KING CHARLS THE SECOND" is what Claude Fable 5.1 pulled out of a cryptogram open since 1899. It took 44 minutes, and the key was the book itself: Urquhart's 32 Proquiritations as a word index.
GPT-6 Astra took the top of Code Arena's WebDev leaderboard from Claude Fable 5.1 on September 5, 1,797 points to 1,762. OpenAI's model holds that lead at $40/Mtoken, which matches the latest Claude pricing.
An internal Anthropic model formalized a complete proof of Fermat's Last Theorem in Lean in 11 days, the last open item on Wiedijk's 100-theorem list. The mathematician funded to do it himself compiled the repo and says it checks out.
84% of Claude's pull requests get merged, against 85% for human-written ones, in a study built on the AIDev dataset. Devin lands at 43% on the same measure, and the paper tracks how those gaps shift over time.
"At max effort it scores 66" on the Artificial Analysis Intelligence Index, the top spot, and the same shop clocks Claude Fable 5.1 at 20% more per task than Fable 5 despite a 75% cache read price cut.
Pay for Claude Fable 5.1's top reasoning level on ARC-AGI-2 and you get 90.0%, the same score the level below already reaches. ARC Prize puts max effort at $4.49 per task in its verified run.
"Web search returns large page-level artifacts, Context7 returns focused documentation slices." In Context7's own benchmark that stripped roughly 99% of Claude Code's fresh input tokens, though cost fell 34.56%.
10 out of 10 alignment benchmarks improved, with no drop in overall performance, in a paper Anthropic published Friday. The automated researcher behind that run costs roughly $4 an hour in API inference against $150 for a human.
30% alone, 99.95% wrapped in an agent harness built with Strands: that's Claude Opus 5 on ARC-AGI-3, per the author's write-up, though NVIDIA reports 100% for the same model inside its own AVO harness.
Unprompted, models name what a fix makes impossible 42% of the time; with the Poka-Yoke skill pack for Claude Code, 81% across 591 blind-graded runs. In return, spotting a raw SQL injection falls from 92% to 69%.
Published Claude Code subagents fail to load at 21.9%, 2.8× the rate of skills, per a census of 445,348 listings. The split traces to Anthropic's frontmatter reference: a subagent's name and description are required, a skill's aren't.
A human expert wrote the prompt, and Claude designed protein binders against 14 out of 15 targets on its own. Between 22% and 35% of them bound successfully, against a typical field rate of 10% to 15%.