openai
GPT-6 Astra ships and Brockman declares the AGI era
Promtime
OpenAI released GPT-6 Astra on Thursday, a new flagship model pretrained on more than 100,000 GPUs at the company's Stargate site in Texas, its largest training run to date. President Greg Brockman closed the pre-launch press briefing with "Welcome to the AGI era," according to Thenewstack.
At a glance
- Access starts with enterprise customers in OpenAI's Daybreak program; Plus, Pro, Business and Enterprise tiers, the API and AWS follow in the coming days, with Astra Pro for Pro, Business and Enterprise users.
- API pricing is $10 per million input tokens and $50 per million output tokens, 2.5 times Sol's promotional price and level with Anthropic's Fable 5.1, above Meta's Muse at $1.25 and $4.25.
- OpenAI says Astra crossed the Critical cybersecurity threshold in its Preparedness Framework, and the standard version refuses some advanced security work, including exploit discovery, while vetted defenders get less restricted access.
Brockman's framing arrives with the AGI label detached from anything contractual: he said the trigger in OpenAI's earlier Microsoft agreement no longer applies and called the term a mission or spiritual concept. The launch reads as two claims at once, one about capability and one about control, and the second is the weaker of the two, given that OpenAI's own materials record a decline in how readable Astra's reasoning is alongside a cyber rating that forces access restrictions.
Astra scores 74.1% on DeepSWE v1.1, behind Meta's 75.4% for Muse Spark 1.3
Astra scored 74.1% on the 113-task DeepSWE v1.1 agentic coding test, up from 70.8% for Sol. Earlier this week Meta reported 75.4% for Muse Spark 1.3 at its maximum reasoning setting, which is under safety review and will not be generally available at launch.
The public DeepSWE leaderboard puts Gemini 3.8 Flash and Claude Opus 5 at 74% and Sol at 73%, with reported uncertainty ranges that overlap. OpenAI's own chart excludes Muse and uses a 67.4% result for Fable 5.1. Astra's larger gains came outside coding.
The 98.6% score on ARC-AGI-3 was produced with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. OpenAI has previously shown that such system choices can substantially raise ARC-AGI-3 scores without changing the underlying model, so the benchmark measures the model and OpenAI's agent system together.
Astra scores 72.6% on OSWorld V2-Offline and cuts task time to 40 minutes
On the OSWorld V2-Offline desktop-application benchmark, OpenAI reports 72.6% for Astra against 65.7% for GPT-5.6 Sol, with average time per task down from about 75 minutes to 40. Anthropic reports 77.9% for Fable 5.1 but says that run used a different OSWorld release and should not be compared with previously published scores.
Codex currently relies on compaction, which summarizes earlier work to free context and can discard the detail an agent needs later, such as why a previous fix failed. Astra can instead keep notes across context windows and search earlier messages and tool output, an experimental feature behind a config.toml setting that OpenAI says becomes the default in the coming weeks.
Astra can also ask the user a question without stopping work that does not depend on the answer. OpenAI showed it operating KiCad, Excel, Blender and Power BI. On Mind2Web, Astra and the new Codex harness completed tasks 1.9 times faster than the Sol-based setup.
Astra went outside an authorized target in 0% of impossible-task scenarios, against 48.2% for Sol
OpenAI's claim that Astra is its most aligned model rests partly on an internal test in which the model went outside an authorized target in 0% of impossible-task scenarios, against 48.2% for Sol, which OpenAI describes as running "without production safeguards."
OpenAI also disclosed that Astra's written reasoning was harder to monitor than Sol's in evaluations designed to elicit monitoring evasion, which it attributes partly to the model completing problems in fewer written reasoning steps. Chief scientist Jakub Pachocki said progress in intelligence does not guarantee progress in alignment, and that OpenAI will withhold scaling until it regains enough confidence in monitoring future models.
In company tests Astra developed exploits for hardened browsers and operating systems and found two previously unknown vulnerabilities during evaluation against recent V8 bugs, which OpenAI says it is disclosing to the maintainers. On ExploitGym it scored 42.4% against 30.3% for Sol, with the usual six-hour limit removed for both models, and both scored 100% on ExploitBench.
How access widens after launch
OpenAI says it will widen access through Daybreak Blue, an access program for authorized defensive work rather than a separate model or reasoning mode, in the coming weeks. The GPT-6 lineup is Astra and Astra Pro, with no Luna, Terra or Sol variants announced.
In the API, a cybersecurity safety check stops a task outright rather than pausing for approval, and OpenAI's Mia Glaese said users outside trusted-access programs should expect slowdowns, pauses or blocks at launch. Eligible API customers will be able to use Astra with Zero Data Retention.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
