GPT-6.1 Sol matches Astra on DeepSWE at a fifth of the cost
Here is the number from the GPT-6.1 Sol launch that people will repeat: at maximum reasoning effort on a scientific benchmark, the new model averages $5.47 per task, while OpenAI's own flagship GPT-6 Astra costs $23.80. More broadly, OpenAI's announcement says Sol nearly matches Astra on agentic coding, computer use and professional work at one-fifth of Astra's standard input and output token prices.
At a glance
- OpenAI shipped GPT-6.1 Sol one week after GPT-6 Sol. API prices stay at $2 and $10 per million input and output tokens, and cached input drops to $0.10 per million.
- On DeepSWE v1.1 it matches Astra at roughly a fifth of the cost and beats GPT-6 Sol's best score by 6.4 percentage points, at lower reasoning effort and cost.
- The benchmark figures are OpenAI's own. Against Anthropic's Sonnet 5.5, sold at the same price, The New Stack calls the results mixed: Sol leads on DeepSWE but trails on AutomationBench.
If you haven't been following, the GPT-6 generation has two versions in this story. Astra is the expensive flagship, and Sol is what The New Stack calls OpenAI's workhorse. According to The New Stack, GPT-6 Sol shipped only a week before this upgrade, on the same day Anthropic launched Opus 5.5, and Sonnet 5.5 followed on Monday at the same $2 and $10 per million tokens as Sol.
On DeepSWE v1.1, Sol matches Astra at roughly a fifth of the cost
DeepSWE v1.1 tests complex software-engineering tasks in real codebases. On it, GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost. It also beats GPT-6 Sol's best score by 6.4 percentage points while running at a lower reasoning effort and cost. The New Stack puts Sol at about 75% on DeepSWE and Sonnet 5.5 at 71%.
GDP.pdf asks models real-world questions about complex PDFs from finance, healthcare, legal and seven other professional fields, including tables, charts, diagrams and fine print. OpenAI says Sol outscores Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings. It also comes close to Astra's state-of-the-art score at roughly one-fifth the cost per task. Reading the chart, The New Stack puts Sol at around 32% and Opus 5.5 at about 29%.
Sol trails Sonnet 5.5 on AutomationBench but costs $0.30 a task against $1.14
AutomationBench 1.0.6 checks whether agents complete end-to-end business workflows using 47 tools across sales, marketing, operations, support, finance and HR. At medium reasoning effort, Sol scores 2.2 percentage points above Opus 5.5 at roughly a third of the cost. That is 4.8 points above GPT-6 Sol at the same setting.
Sonnet 5.5 does better here. The New Stack has Sol at around 36% and Sonnet 5.5 at 44.7%, though Sol costs $0.30 per task against $1.14. A footnote from OpenAI says the datapoint for Claude Fable 5.1 understates its real cost. It leaves out fallbacks, which happened on about 40% of tasks.
On the offline set of OSWorld 2.0, agents work through long computer-use tasks, both everyday and professional. At maximum reasoning effort Sol beats GPT-6 Sol by seven percentage points for less than half the cost. It lands within 2.1 points of Astra at maximum effort for roughly one-seventh the cost per task. OpenAI reports partial reward from the v2026.08.08 release.
Terminal-Bench Science costs $5.47 per task on Sol and $23.80 on Astra
Terminal-Bench Science 0.1 covers data analysis, simulation and theorem proving. At maximum reasoning effort, Sol more than doubles GPT-6 Sol's score at less than half the cost per task. It averages $5.47 per task, against $23.21 for Opus 5.5 and $23.80 for Astra. OpenAI calls that over 75% cheaper than either.
Astra still has the top score among the tested models at 68.1%, and OpenAI says it should be used for the hardest scientific research. The New Stack goes the other way and argues that these results make Astra hard to justify for most use cases.
Sol also gets facts wrong less often. At low reasoning effort, the share of responses with at least one factual error falls from 11.4% on GPT-6 Sol to 7.7%, a reduction of roughly 32% and Sol's biggest factuality gain. Across the tested reasoning settings, Sol's error rate stays within 1.9 percentage points of Astra's at less than one-fifth the cost per task.
Sol made no attempt to bypass an automated safety reviewer in OpenAI's tests
OpenAI reports large gains over GPT-6 Sol in its alignment evaluations, which bring Sol closer to Astra. In hard test cases, the new model fails less often than GPT-6 Sol on three things: admitting when a search tool is broken, following explicit restrictions, and avoiding unauthorized outcomes during agentic tasks. The New Stack says Sol fails to report a broken search tool only 2.8% of the time, rather than guessing.
Like Astra and GPT-6 Sol, it made no attempt to bypass an automated safety reviewer. OpenAI stresses that these evaluations deliberately test difficult situations and do not measure failure rates in typical use. Full details are in the GPT-6.1 Sol system card addendum.
Cached input costs $0.10 per million tokens, 95% below the standard rate
Standard API prices are $2 per million input tokens and $10 per million output tokens, unchanged from GPT-6 Sol according to The New Stack. Cached input costs $0.10 per million, 95% less than standard input and 50% less than GPT-6 Sol's cached rate.
Caching matters for agents because they send the same context over and over: the system prompt, the tool definitions, the conversation so far. When the start of a request repeats something the provider has already processed, those tokens are billed at the cached rate, not the full one. It works like a print shop that charges full price to set up a page once and a fraction for each reprint.
The model is live for Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, and in the API as gpt-6.1-sol. It is not yet available in Chat. On Twitter, OpenAI called Sol the most cost-efficient model for its performance available today.
Almost every number here comes from benchmarks OpenAI ran and charted itself. By OpenAI's own description, the factuality and alignment tests are deliberately hard cases that do not reflect typical use, and the Claude Fable 5.1 footnote shows that fallback costs are not counted the same way for every rival. In our view, Sol's case against Sonnet 5.5 rests on cost: $0.30 per AutomationBench task versus $1.14, for a lower score.
When Ultrafast reaches Codex. OpenAI says GPT-6.1 Sol Ultrafast will arrive in the coming days, with token generation up to 8x faster than standard speed in Codex. No Ultrafast price has been given, and there is no date for Sol in Chat. The New Stack found only a few benchmarks with results for both Sol and Sonnet 5.5, so a fuller comparison is still missing.
Related stories
- OpenAI's half-price GPT-6 Sol and Luna aren't a promo
- GPT-6 Astra lands on the API at $10 in, $50 out
- OpenAI's $200 Pro plan comes back worth half the API spend
- OpenAI shelves its next model after internal safety tests
- Code strings point to a $50 production plan for OpenAI's API
- GPT-6 caching update cuts cached input costs by up to 90%
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
