openai
Astra's benchmark numbers changed after the blog went live
Promtime
openaiOpenAI changed several evaluation figures in its GPT-6 Astra announcement after the post went live on September 3, including a hallucination rate that dropped from 4.2% to 2% and later returned to 4.2%, according to Fortune, which compared internet archive snapshots of the page.
At a glance
- Internet archive snapshots show the hallucination figure holding at 4.2% through 3:11 p.m. ET, then changing in the 5:20 p.m. capture along with four other metrics.
- On FrontierMath Tier 4 (v2), Anthropic's Fable 5.1 fell from 87.8% to 78% and now reads 83%, while GPT-5.6 Sol moved from 83% to 80.5% and back to 83%.
- OpenAI says most evaluations carry noise of a few percentage points depending on checkpoint, scaffold and run, and that the launch-blog fixes were meant to reflect its best estimate of model performance.
The episode reads less like one company's lapse than a structural weakness in how the industry keeps score. Benchmark tables are marketing surfaces as much as measurements, and when the same figure moves twice in a day, the value of comparing vendors on published numbers erodes for everyone reading them. Ahead of a possible 2027 IPO, confusion over its own metrics likely works against the clean narrative OpenAI would prefer to present.
Five metrics changed in the 5:20 p.m. snapshot of the blog post
The first internet archive snapshot of the announcement, taken at 2:23 p.m. ET on September 3, put Astra's hallucination rate at 4.2%. It held through four more captures, the last at 3:11 p.m., then appeared as 2% in the sixth snapshot at 5:20 p.m., alongside four other changed metrics.
GPT-5.6 Sol's hallucination rate moved from 12.2% to 9.4% in the same edit, and both figures have since returned to their original values. Astra's coding score also rose in the later versions, from 57.7% to 57.9%, and Sol's result on OpenAI's internal version of the ExploitBench cybersecurity evaluation went from 5.5% to 11.5%.
Two changes favored Anthropic. On HealthBench Professional, Claude Fable 5.1 improved from 56.6% to 58.1% and Opus 5 from 54.5% to 56.4%. Scores for rival models in the post are usually taken from published leaderboards rather than produced by OpenAI running its own assessments on competitors.
OpenAI published the post shortly after 2 p.m., then retracted it
OpenAI had scheduled the announcement for 2 p.m. ET and published it shortly after that, then pulled it. OpenAI said it could not disclose the reason but that it was unrelated to the benchmark figures, having first cited a bug in its content management system and then an internet outage.
The OpenAI account on X posted the link at 3:32 p.m., and it returned an error message. At 3:50 p.m., Sam Altman shared it with the line: "We hit a little snag getting the blog post deployed, but it is really great." Commenters and Fortune still saw the error, and the page loaded properly about an hour later.
A disclaimer on the blog reads "Evaluation scores are the maximum at any effort," with further caveats in footnotes on each metric. Different research teams at OpenAI oversee different metrics and are responsible for calculating them and reporting them to a central team for publication.
Astra's ARC-AGI-3 score went from 98.6% in the draft to 99.99% live
The edits began before publication. An embargoed pre-publication draft sent to Fortune and other media organizations listed Astra's ARC-AGI-3 score as 98.6%, while the live blog says 99.99%. A spokesperson said the company always verifies evals before publication, so adjustments between draft and final version are normal.
The Arc Prize Foundation, which created the benchmark, measured Astra at 99.9% in an independent assessment when the model was given a particularly powerful harness, and at 63% with the benchmark's standard harness, ahead of any other AI model in public release. OpenAI said harness, reasoning level and other factors inform evals.
Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab, said re-running evaluations under different conditions can be done in a tight timeframe and is better for marketing. They said the GPT-6 Astra system card offers "barely any details" about the internal hallucination benchmark and does not include the number of test items.
Whether Sol's ExploitBench score reverts
OpenAI said it is investigating reverting Sol's ExploitBench figure to 5.5%, because the 11.5% result reflects a reasoning level that is not commercially available for that model. Vincent Sunn Chen, who leads benchmark and evaluation research at Snorkel AI, said he would like industry norms requiring companies to state what changed in the measurement setup whenever they revise published scores.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
