Skip to content

openai

OpenAI's math results cost about three hours of Pro each

Promtime

On average, each new mathematical result in OpenAI's latest release took about as much compute as three hours of ChatGPT Pro thinking. The figure comes from OpenAI's own announcement, which publishes a broad range of results produced by an internal frontier model that OpenAI has not released.

At a glance

  • The results are published in a GitHub repository with protocols for paper revisions and citations, following advice from the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study.
  • According to Testing Catalog, the model was posed approximately 4,000 problems; separately, OpenAI says many of its proofs come with Lean formalizations that a computer can check.
  • The catch is that the announcement mentions statistics on attempted problems but never says how many of those approximately 4,000 were solved, and the model itself is still unreleased.

OpenAI measures the compute in hours of ChatGPT Pro thinking

OpenAI calls the source of the results an internal frontier model. It is unreleased, and the announcement does not name it. Testing Catalog quoted the repository as saying the vast majority of results were obtained with the same procedure using that model, and that over the course of the evaluation it was posed approximately 4,000 problems.

The compute is given in a unit any ChatGPT subscriber will recognise. OpenAI estimates the spend in terms of Pro usage on ChatGPT and says the average result used the equivalent of roughly three hours of ChatGPT Pro thinking. Testing Catalog reports the same three-hour average per result.

The repository also includes statistics about the number of attempted problems and 10 summaries of the model's reasoning. OpenAI says it is publishing these details about how the results were obtained to promote scientific transparency and openness.

OpenAI shaped the release on advice from the Institute for Advanced Study

OpenAI says it has been consulting the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on best practices for sharing results with mathematicians. It says it drew on the group's advice and public recommendations for this release. The results are at github.com/openai/math, along with protocols for paper revisions and citations.

OpenAI does not present GitHub as the permanent home for the results. It says it is still exploring community-hosted alternatives that meet the committee's guidelines. For future releases, OpenAI says it will improve the papers themselves: better citations, clearer mathematical exposition and a presentation of the results that is easier to follow.

OpenAI plans to fund workshops and conferences on AI-produced results

OpenAI says it will fund a series of workshops, conferences and special programs on understanding major results produced by AI. It says it wants this progress to push the frontier of human knowledge and enable further progress in mathematics, and to give scientists state-of-the-art capabilities directly.

The company also plans to keep evaluating its internal frontier models on mathematics and other sciences so it can speed up building tools for those fields. On the model itself, OpenAI says it is working to release it responsibly. It adds that it will keep acting on community feedback and update its standards for future disclosures of major scientific advancements.

Many of the proofs come with Lean formalizations a computer can check

Lean is a programming language for writing a mathematical proof precisely enough for a computer to verify it. A proof written for humans skips steps that experts consider obvious. A Lean proof spells out every step, and the checker rejects the whole proof if any link in the chain fails.

Think of it as a spell-checker for logic. It cannot tell you whether a result is interesting, only whether every step is valid. With a formalized proof, you do not have to trust the model or a referee's patience. OpenAI says the repository covers many of the proofs, not all of them, and it will add more formalizations as it obtains them.

Oddly, the announcement never says how many of the approximately 4,000 problems cited by Testing Catalog the model actually solved, even though OpenAI says the repository includes statistics on attempted problems. Without that ratio, the three-hour figure reads as the average cost of a published result rather than of the whole search, so it may understate the total compute spent.

When the model reaches researchers

No release date has been given for the model behind these results. OpenAI promises to share more about the workshops, conferences and special programs only in the near future, also without dates. Two open questions remain for the next update: whether OpenAI states a solve rate for the approximately 4,000 problems, and whether the results move to a community-hosted home that meets the committee's guidelines.

Related stories

  1. Mathematicians recall an OpenAI promise it isn't aware of
  2. Mathematicians get a say in how OpenAI reports its math
  3. OpenAI trains GPT-6 Astra on real Ironclad contract work
  4. Blocked from the web, an OpenAI agent tunneled out via DNS
  5. OpenAI wants a safety case before every frontier RL run
  6. OpenAI's swarm spent days fighting a check it only inferred

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.