Skip to content

openai

OpenAI trains GPT-6 Astra on real Ironclad contract work

Promtime

On one contracting task, GPT-6 Astra met about 94% of the grading criteria in an estimated 20 minutes. GPT-5.6 Sol met about 85% in an estimated 32. The comparison comes from OpenAI's write-up of its research partnership with Ironclad, the AI contracting company that helped turn real legal-operations work into training and evaluation tasks.

At a glance

  • OpenAI is working directly with a small number of software companies to turn hard workflows into research problems, and Ironclad is the first, supplying tasks, success criteria and hosted product environments.
  • Across 11 research tasks, GPT-6 Astra at Max reasoning averaged 55.0% against 41.6% for GPT-5.6 Sol at High reasoning, and estimated time per attempt fell from 37.0 to 19.2 minutes.
  • All of these numbers come from OpenAI's own research evaluation, and an internal model built during Astra's development scored 63.7%, so Astra is not the top scorer on these tasks.

If you haven't been following: when OpenAI introduced GPT-6 Astra, it showed the model doing professional work on a computer, from preparing documents to testing websites. OpenAI says the next goal is agents that handle specialized business software. That means understanding a company's business rules, running multi-step workflows and checking that the finished work meets the original requirements.

Eleven tasks, each graded against 8 to 50 criteria

Ironclad employees and people at OpenAI who use Ironclad helped the researchers pick 11 tasks across legal, commercial and procurement work. Examples include setting up nondisclosure agreements, building procurement approval processes, and updating a reusable legal clause so it reflects the jurisdiction a requester selects. OpenAI estimates that an experienced user would need about 30 to 40 minutes per task on average.

Each task is scored against 8 to 50 criteria, depending on how complex it is. OpenAI says this shows researchers which parts of a task a model got right and where it fell short, so a task is not graded as a single pass or fail. Ironclad's team, according to the post, was instrumental in defining what success looks like.

OpenAI presents the partnership as a way for Ironclad to bring its customers' problems into the development of the underlying model. A contracting process has to handle exceptions without breaking the rules a business depends on. As the models improve, OpenAI writes, Ironclad could use them in more of its own products.

Astra scored 55.0% where GPT-5.6 Sol scored 41.6%

GPT-6 Astra is OpenAI's first frontier model trained on Ironclad tasks. OpenAI tested each model at the setting where it scored highest: Max reasoning for Astra, High reasoning for GPT-5.6 Sol. Across the 11 research tasks, Astra averaged 55.0% and Sol 41.6%, which OpenAI describes as a 32% higher average score.

Estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra, or 48% less. The single-task clips show the same pattern: about 94% of criteria in an estimated 20 minutes for Astra against about 85% in 32 for Sol. An internal model used during Astra's development reached 63.7% on the 11 tasks, and OpenAI says it aims to bring those gains to future models.

The agent has to test its own approval rules on both sides of the threshold

Take a legal operations team setting up a process for buying software. Finance may need to approve purchases above a certain amount, Security may need to review certain requests, and Legal may need to review nonstandard terms. Someone has to turn that short list into an intake form, document templates, approval rules and a record of the final agreement.

An agent doing this job has to keep every requirement in view as it clicks through the software. It must configure Finance approval above the spending threshold, then check that requests above and below it follow the right paths. Picture a plumber who fits every pipe correctly but never turns the water on. Until the whole system runs, the job is not finished.

The training setup has three parts. People who know the work help choose the tasks, every task comes with detailed criteria, and Ironclad hosts environments of its product where models can practice. On top of that, researchers built synthetic training tasks around representative workflows and used reinforcement learning, so the models improve through repeated attempts and feedback on what went wrong.

Every figure here comes from OpenAI's own evaluation of 11 tasks. The post also concedes that human oversight still matters and that a full contracting platform remains essential, because an agent that loses track of one rule midway limits what it can be trusted with. Oddly, the post calls the times both estimated and simulated, but its main text never explains how those minutes are calculated.

Who OpenAI wants as the next partner

OpenAI is inviting a small number of software companies to work with its research and engineering teams on tasks that today's agents still cannot reliably complete. It wants a concrete example, evidence of where the agent fails and a way to judge success. Partners also need people who know the work well, a secure test environment and data that is safe to use for research. OpenAI has not named a second partner or said when the 63.7% model's gains will reach a released model.

Related stories

  1. Three computer-use benchmarks fall to GPT-6 Astra
  2. ChatGPT Auto-review goes free for ChatGPT sign-ins
  3. OpenClaw agents kept reading memory after it was disabled
  4. Wikimedia finds OpenAI agents in its sandboxes and Etherpad
  5. OpenAI's agent audit costs over $500,000 a day
  6. OpenAI's rogue-agent warnings reach more than 100 groups

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.