openai
OpenAI denies anyone looked at user data for Navier-Stokes
Promtime
openai"Did any human or agent look at user data as part of the Navier Stokes effort? No." That is one half of OpenAI's answer to the mathematicians who accused it of being dishonest about training data, posted on Twitter; the other half is a yes, because the company says it uses user feedback and de-identified data to improve ChatGPT and Codex, "and so does every LLM company."
At a glance
- OpenAI says no human and no agent saw user data during the Navier-Stokes effort, answering mathematicians who accused the company of being dishonest about its training material.
- The same post confirms the ordinary pipeline: user feedback and de-identified data feed improvements to ChatGPT and Codex, and OpenAI's own pages say conversations keep training the model unless you opt out.
- Nowhere does the statement name the material that actually went into the Navier-Stokes work, and the denial is scoped to human and agent access rather than to the training pools themselves.
If you have not been following: according to Nature, OpenAI announced on 8 September that it had generated a solution to the Navier-Stokes equations showing they can break down, cracking one of the Clay Mathematics Institute's Millennium Prize Problems, which come with a million-dollar prize. TU Dresden geometry professor Andreas Thom then wrote on Mastodon that his own ChatGPT interactions may have fed the announced results.
OpenAI separates access to user data from training on it
The post pulls apart two questions that critics had been folding together: who or what could look at user data during the Navier-Stokes effort, and what the training pipeline does in general. On the first, OpenAI's answer is that no human and no agent did. On the second: "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."
That is not the first scoped denial in this story. According to Aichatdaily, OpenAI's Navier-Stokes announcement said its researchers and agents did not see any of the other team's work through any means until it was released publicly, and that no specific user data was used.
Thom's complaint started with non-sofic groups in August
The row did not begin with fluid dynamics. According to Aichatdaily, Thom works on non-sofic groups, roughly infinite mathematical structures that cannot be approximated by finite ones, and that was one of ten results OpenAI announced in August with what he described as great fanfare.
That announcement drew criticism in mathematical circles for failing to credit recent work by Thom and Gábor Kun, and the company quietly amended its writeup afterwards, Aichatdaily reports. What struck Thom most, he said, was OpenAI's detailed command of techniques that were neither the most obvious nor the most promising routes to a solution at the time.
According to Aichatdaily, he emailed OpenAI researchers Sébastien Bubeck and Mark Sellke, the latter also a statistician at Harvard, to ask whether his ChatGPT conversations were part of the training data or accessible to the reasoning process; the reply addressed only whether the chatbot could directly retrieve his conversations.
Rival zero-viscosity papers landed on 7 September, OpenAI's claim on 8 September
According to Nature, OpenAI had been testing its latest prototype on all six unsolved Millennium Problems and decided on 1 September to put resources behind Navier-Stokes, after hearing rumours that Levent Alpöge of Harvard and Tristan Buckmaster of NYU had solved a version of it with Anthropic models. Buckmaster has publicly questioned whether OpenAI's models benefited from his own use of Codex while he worked on the same problem, according to Aichatdaily.
On 7 September, Alpöge and Buckmaster released a paper claiming a solution for the simplified zero-viscosity case, in which the fluid reaches infinite speed, using Anthropic's Claude together with OpenAI's Codex and Astra. The same day, Anima Anandkumar of Caltech and collaborators released their own zero-viscosity solution, built with a physics-informed neural network rather than a general-purpose large language model.
De-identified data still starts as real conversations
De-identified means the text has been stripped of what ties it to a person. On its own pages OpenAI says ChatGPT improves by further training on the conversations people have with it, unless users opt out, and that it has built technologies to help models learn useful general patterns rather than private information about individuals.
Think of an instructor working through a stack of anonymised exam papers: what carries into next year's teaching is the shape of a good argument, not whose handwriting it was. The line OpenAI draws sits there, between general patterns and the content of any single chat.
What the statement does not do is name the material behind the Navier-Stokes work. The denial covers human and agent access; the yes covers a pipeline that, by OpenAI's own description, keeps learning from conversations by default. In our view the mismatch in scope is the odd part: a question about training pools has now been answered twice with an answer about who or what could look something up.
The general-case paper
According to Nature, Alpöge and Buckmaster said a solution to the more general problem would be released soon, and no date has been given for it. The next thing to watch is whether the zero-viscosity papers and OpenAI's claim survive the reading mathematicians are giving them now, and whether any lab spells out, in terms a researcher can check, what separates a training pool from a specific user's chats.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
