openai
Reasoning off, Astra still hits 96.7% on ARC-AGI-3
Promtime
openaiARC Prize’s published table put GPT-6 Astra at 62.7% on ARC-AGI-3 under its standard harness, while OpenAI’s Provider Adapter recorded 99.9% for the same model. Thenextweb reported that the two results used different software around the model, rather than different model weights.
At a glance
- ARC Prize’s standard setup gives models a minimal common interface and leaves them to decide which notes survive between requests, while OpenAI’s adapter preserves opaque reasoning state and compacts long conversations.
- With reasoning effort set to none, Astra reportedly reached 96.7% inside OpenAI’s adapter, exceeding its 62.7% result at maximum reasoning in the standard harness by 34 percentage points.
- The widely circulated comparison paired Astra’s adapter result with GPT-5.6 Sol’s 7.8% standard-harness result, although the like-for-like comparison is Astra at 62.7% against Sol at 7.8%.
The results show that benchmark scores increasingly describe an assembled system, not only a model’s underlying weights. That distinction appears especially consequential when a benchmark result is used to support a claim about general intelligence. It also complicates independent comparison: a provider can expose documented tools through an API, while outsiders may still be unable to reproduce the same state handling and context management.
OpenAI’s adapter delivered 99.9% while the standard harness delivered 62.7%.
ARC Prize reportedly ran Astra at maximum reasoning in its standard harness for $26,098. In OpenAI’s Provider Adapter, a high-reasoning run scored 99.9% and cost $18,817. Both harnesses solved 167 game-reasoning pairs; on that shared subset, ARC Prize measured the adapter runs using 49% fewer tokens and completing about 3.66 times faster.
ARC Prize said it was not claiming Astra was AGI, and co-founder Mike Knoop wrote that the available evidence did not justify that label. The foundation plans to publish results from both harnesses side by side. François Chollet separately gave the standard-harness result as 66% in a post, while ARC Prize’s published table lists 62.7%.
Five launch metrics reportedly changed after OpenAI published its post.
Fortune compared archived versions of OpenAI’s launch post and reported that Astra’s hallucination rate first appeared as 4.2%, changed to 2% by 5:20pm, and later returned to 4.2%. The embargoed media draft reportedly listed 98.6% for ARC-AGI-3, whereas the live post displayed 99.99%.
Fortune also reported that Anthropic’s Fable 5.1 FrontierMath score moved from 87.8% to 78%, then settled at 83%. Sol’s ExploitBench figure changed from 5.5% to 11.5%; OpenAI said it was investigating a reversion because 11.5% used a reasoning level Sol does not offer commercially. OpenAI reportedly attributed ordinary variation to checkpoint, scaffold and evaluation-run changes.
Independent testing placed Astra at 67 on a coding-agent index.
Artificial Analysis reported an Astra score of 67 on its Coding Agent Index in Codex, level with Claude Opus 5 and Fable 5, while Fable 5.1 scored 70. Its Intelligence Index put Astra at 61, equal to the model it replaces, five points below Fable 5.1 and below Meta’s Muse Spark 1.3.
Artificial Analysis listed Astra at $10 per million input tokens and $50 per million output tokens. It reported that maximum-effort tasks cost 75% more than Sol’s. Astra’s knowledge-benchmark hallucination result reportedly fell from 92% to 51%, and long-horizon knowledge work gained about 80 Elo, while an economically weighted task benchmark and several named task categories regressed.
Private tests after the revisions Artificial Analysis rebuilt its index as version 4.2 the following day, adding harder tasks and more private test sets that it says are intended to limit gaming. Separately, Stanford researchers Anka Reuel and Mike Hardy described repeated evaluation changes intended to improve a result as “benchmaxxing” and said Astra’s system card offered barely any detail about its internal hallucination evaluation, including no test-item count.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
