openai

Three computer-use benchmarks fall to GPT-6 Astra

Promtime

openai

OpenAI says its new flagship holds state-of-the-art results on Agents' Last Exam, AutomationBench and ScreenSpot Pro, benchmarks that measure computer workflow tasks across professions. The claim went out on X without any accompanying numbers.

A second post stretches the list to FrontierMath Tier 4, ARC-AGI 3 and TerminalBench-4.0, and calls Astra a major advance for scientific discovery with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro. On ARC-AGI 3, the run ARC Prize published on September 3 put Astra at 99.9% under the Provider Adapter harness.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.