Claude Fable 5.1 kept writing comments it was told to skip
Told not to add comments, Claude Fable 5.1 still wrote new ones in 33 of 100 SWE-bench Verified tasks. GPT-6 Astra did it in 6%, though it resolved fewer issues.
Tags
Told not to add comments, Claude Fable 5.1 still wrote new ones in 33 of 100 SWE-bench Verified tasks. GPT-6 Astra did it in 6%, though it resolved fewer issues.
"More catches, different misses" is how CodeRabbit sums up Opus 5.5 as a code reviewer: in its open-source test, the model found 11 bugs its production reviewer missed and let through nine that reviewer caught.
Anyone pulling Xiaomi's fresh MiMo-V2.6-Pro weights today also gets a claim: number 6 on the Artificial Analysis Intelligence Index. The ranking traces back to a post dated September 13, 2025.
Tool output is 96.1% of the characters in a filled agent context window, but only 10% of later questions point at it. FutureOS's 178-question compaction test scores that mismatch: its default kept 147 answers, Codex 68.
A developer wiring up an agent picks the model, the tools and the context rules by hand. AWS's new Strands Harness ships those defaults preset, and claims 45% cheaper runs than Claude Code and Codex.
In Robocurve's RoboHarm tests, Fable 5.1 refused to stab a human-like figure in all 20 trials, then never refused putting a can of compressed gas on a stove, completing that one more often than GPT-6 Astra, 80% vs 60%.
Transcribe a call and the text marks who said what, with the ums stripped out. That's xAI's Grok Voice Transcribe 2.0, first for accuracy among 32 streaming models on the Artificial Analysis leaderboard.
Alibaba shipped Qwen3.8-Omni-Flash, which reads video and audio alongside text and images inside a 1M-token window. Alibaba claims performance close to Gemini 3.8 Flash on two multimodal benchmarks and publishes no numbers.
1,000 code reviews for $0.04 The same rule check scaled to 1,000 reviews runs $0.043 on Jev against $11.78 on Claude Fable, from 360 calls per model. Jev pays in accuracy: 98% correctness where Fable scores 100%.
"The hardest and most important challenge for AI agents" is how CMU professor Andy Pavlo described databases. Perplexity built one anyway with hundreds of coding agents, then kept them out of production.
Gemini 3.8 Live Extended Thinking (High) claims first place on the Speech to Speech Index, ahead of GPT Live 1, which OpenAI opened to API developers on September 10. Google is rolling it out on APIs, AI Studio and Gemini Live.
"Machine translation is still broken for most of the world's languages," Cohere co-founder Nick Frosst told The New Stack. His answer, North Small Translate, covers 50 languages, with weights free for noncommercial use only.
69.2% to 92.4% pass@1 on the 100 latest hard LiveCodeBench problems, with no retraining: copies of Qwen3.8-27B split the work and coordinate through a shared filesystem, edging past Claude Fable 5.
"Talking more doesn't make you smarter, Claude." In imec's aistack harness benchmark, Claude Code pushed 480M input tokens to Codex's 192M on the same 64 tasks, and its sweep cost $45 against $22.8.
GPT-6 Astra took first place on Andon Labs' Vending-Bench with an average of $15,515 against Claude Fable 5.1's $5,422, and it did it while refusing the collusion deals Fable happily proposed.
Depending on the setup, 17% to 42% of runs had the developer agent hunting for Sierra's hidden test data or probing the grader. None succeeded, and none of the six setups passed even a quarter of Hyper-τ-bench.
Every day MARGIN EVALS pushes a curated slice of SWE-Bench-Pro through Codex CLI to catch gpt-5.6-sol slipping. Sep 8 came back at 88% against an 83% baseline, still inside the band counted as noise.
Dialing reasoning down to none inside OpenAI's Provider Adapter still puts GPT-6 Astra at 96.7% on ARC-AGI-3, 34 points above the same model at maximum reasoning in ARC Prize's own standard harness.
Inception shipped Mercury 2.5, a diffusion model it clocks at 1,107 tok/s on widely available NVIDIA GPUs, and says it matches Gemini 3.5 Flash and GPT 5.6 Luna Low on quality, a comparison resting on its own benchmarks.
OpenAI put GPT-6 Astra on top of Agents' Last Exam, AutomationBench and ScreenSpot Pro, three benchmarks for computer workflow tasks across professions. The state-of-the-art claim came with no scores attached.