OpenAI says its new flagship holds state-of-the-art results on Agents' Last Exam, AutomationBench and ScreenSpot Pro, benchmarks that measure computer workflow tasks across professions. The claim went out on X without any accompanying numbers.
A second post stretches the list to FrontierMath Tier 4, ARC-AGI 3 and TerminalBench-4.0, and calls Astra a major advance for scientific discovery with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro. On ARC-AGI 3, the run ARC Prize published on September 3 put Astra at 99.9% under the Provider Adapter harness.

