openai
OpenAI walks out of Caltech's mathathon over "slop" letter
Promtime
openaiMathematicians at Caltech have a name for the genre of AI proof announcement that lands on social media and leaves the checking to everyone else: "slop mathematics." A day after they said so in an open letter, OpenAI stopped sponsoring Caltech's math hackathon, Business Insider reports.
At a glance
- OpenAI research lead Dan Roberts said on X on Thursday that the company has ended its sponsorship of the event, and that the rapid progress of AI in mathematics is disruptive.
- Round one of the mathathon gives each team 40 hours and $20,000 in tokens against an open problem; teams with promising, well-explained answers get six months and more tokens in round two.
- OpenAI and Anthropic had pledged $2 million in AI credits between them; per Business Insider, how much of that came from OpenAI, and what happens to those tokens now, was unclear on Thursday.
If you have not been following the run-up: according to Quanta Magazine, OpenAI said in May that an internal model had produced a counterexample to the unit distance problem, a conjecture Paul Erdős posed in 1946, the first historically significant proof from an AI model, though human mathematicians substantially improved on it within weeks. On August 1, again per Quanta, an unreleased model called Astra was credited with ten more advances, including three further Erdős problems.
The letter's term for the genre is "slop mathematics"
The letter went out on Wednesday, signed by current and former Caltech mathematicians, and it says the event "is likely to have destructive impacts for the mathematical community." What they object to is a sequence: an AI company claims an outstanding theoretical problem, then leaves the proofs and the practical work to researchers.
After the latest result is dropped, research mathematicians are compelled to step in to properly verify, disseminate, and sometimes discredit entirely the claimed results. This labor goes uncompensated, uncredited, and unacknowledged.
They also frame the sponsored hackathon as advertising, saying the companies "will take credit for the effort of talented undergrads." According to Proofs and Prompts, which carries the letter, it describes an arms race to be first to prove prominent open conjectures, with new AI-generated results appearing weekly, and calls the practice research misconduct.
OpenAI said on Tuesday its agents had solved Navier-Stokes
Tuesday is when the temperature spiked. OpenAI announced that its agents had solved the Navier-Stokes equations, a 90-year-old set of formulas that predict how liquids and gases move. Roberts, replying on Thursday, said OpenAI wants to engage with the math community more on the best way to integrate the technology and communicate its impacts.
According to The Verge, Navier-Stokes is one of the seven Millennium Prize Problems, each carrying a $1 million reward, and OpenAI used an internal model it described as more powerful than the newly released GPT-6 Astra together with 10,000 concurrent agents. The Verge reports OpenAI began training that internal model on August 28 and does not plan to take the prize money.
Tristan Buckmaster published on the same route one day earlier
Buckmaster, a mathematician at NYU, and another academic published research on the problem on Monday. On Tuesday he said OpenAI might have used his conversations with AI chatbots, including OpenAI's, to arrive at its solution. OpenAI said it could not rule out that "de-identified data derived from their usage of our products helped improve our models."
According to The Verge, Buckmaster's co-author is Anthropic researcher Levent Alpöge, and the pair worked with OpenAI's Codex and Anthropic's Claude; OpenAI also said no specific user data was accessed. The Verge quotes OpenAI technical staff member Sébastien Bubeck saying the team saw none of their work until it was public, and Buckmaster answering on Mastodon that OpenAI was admitting to training data from after his result.
The mathathon is billed as practice, and it is run by undergraduates
The application page describes the event as a practice ground for young mathematicians to "responsibly use AI tools to augment human understanding of mathematics." First round: 40 hours and $20,000 in tokens per team on an open problem. Second round: six months and additional tokens for the teams whose answers are judged promising and well explained.
The organizers, many of them undergraduates, replied on Thursday. They wrote that events like the mathathon encourage young people to stay excited about mathematics and engage their interests with modern tools, that it is their first time running an event at this scale, and that they are grateful for the feedback.
Final-answer benchmarks do not look at the proof
Here is the mechanical reason a claimed result becomes someone else's job. A model's answer to an open problem is a document, not a verdict; somebody has to read the argument line by line and decide whether each step holds, and that reading is the uncompensated labor the letter describes. It is the difference between checking the answer at the back of the textbook and marking the working.
A paper introducing the Open Proof Corpus, a dataset of more than 5,000 human-evaluated LLM-generated proofs, makes the same point about measurement: on arXiv the authors write that standard final-answer benchmarks such as AIME and HMMT fail to capture the full range of mathematical capabilities, because they do not require models to produce proofs or detailed intermediate steps.
The money is the part nobody has spelled out. Business Insider reported that the split of the $2 million pledge between the two sponsors was not disclosed, and that the fate of OpenAI's share was unclear on Thursday. Oddly, the withdrawal leaves the design the letter objected to untouched: the same open problems, the same 40-hour clock, now with one funder fewer.
The October 30 clock
According to Proofs and Prompts, the letter says the mathathon was scheduled to begin hosting on October 30, 2026. Anthropic remains a sponsor. What is not stated anywhere in the material: whether OpenAI's portion of the credits gets replaced, whether the token allowance per team changes, or whether the six-month second round proceeds as designed. The organizers have said only that they are still listening to feedback.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
