The mathematician, the chatbot, and the clause nobody reads

Two researchers spend a year proving a Millennium Problem, uploading every draft of their work into Codex. Days before publication, OpenAI presents an overlapping solution obtained with an internal model. Asked about it, the company cannot rule out that the two mathematicians' sessions contributed. It isn't theft — it's the clause almost anyone using a consumer AI product has probably accepted without realising it.
Tristan Buckmaster (NYU) and Levent Alpöge (Harvard) spent nearly a year attacking one of the finite-time blow-up problems in the Navier-Stokes family — forced 3D Euler and Boussinesq equations — one of the seven Millennium Problems set by the Clay Mathematics Institute, carrying a one-million-dollar prize. They worked with several AI models in parallel: Claude, Codex, GPT-5.6 Sol and Astra. By their own account, every draft of the project passed through Codex sessions.
By 15 August 2026 they had blow-up results for Boussinesq and 3D Euler; by 22 August, a proof verified in Lean. Before the work could be published, OpenAI presented ten "advances in mathematics and theoretical computer science," including a solution to the same problem, attributed to an internal version of its Astra model.
Asked whether the two mathematicians' Codex sessions might have contributed to training the model behind the competing result, OpenAI responded — according to Fortune and VentureBeat — that it could not rule out that "de-identified data derived from usage of its products" had contributed to improving its models.
As of 1 September 2026, none of OpenAI's ten results has cleared traditional academic peer review, and a dispute over credit is underway between the parties. This article does not take a position on the underlying mathematical priority dispute — the interesting part, for anyone working in AI governance, starts here.
Why it isn't (technically) theft
Consumer Codex sessions are training data by default.This is not a bug, not unauthorised access: it is the declared business model in the terms of service. Anyone using Codex — or any equivalent consumer product — without enabling opt-out, or without an enterprise plan that contractually excludes it, has agreed, by clicking "accept," that their conversations may contribute to improving the models.
Buckmaster did not (necessarily) suffer a data breach. He discovered, in the hardest possible way, the difference between two sentences that sound identical to anyone who doesn't read the fine print:
- "We don't train models on your data" — true for enterprise/API plans with explicit no-training clauses, often at extra cost.
- "You may have consented to becoming training data" — true for most consumer accounts, free or pro, absent an explicit opt-out.
These are two different products, two different contracts, two different risk levels — sold under the same brand, often behind an identical interface.
The problem no external check can resolve
The most uncomfortable part of this case isn't whether OpenAI "stole" an idea. It's that there is no independent technical way to verify it, either for Buckmaster or, in any evidentiary sense, for OpenAI itself. A language model keeps no readable log of "which document influenced which output": the evidence is statistical, circumstantial, never conclusive in the strict sense. Even the most well-intentioned company cannot, technically, produce a negative proof ("we didn't use that data") that would withstand outside scrutiny.
This is exactly the kind of risk a compliance/AI governance function needs to handle before signing a contract with an AI vendor, not after discovering a problem:
- Read the training clause, not the marketing. "Enterprise-grade" and "we don't train on your data" are not automatic synonyms: they need to be checked line by line in the DPA, not on the product page.
- Separate plans by intended use. If drafts, prototypes, contracts, expert reports, or not-yet-filed intellectual property circulate inside your organisation, consumer/free access is probably the wrong plan — no matter how convenient it is.
- Treat the absence of proof as residual risk, not reassurance. If a vendor cannot technically demonstrate it hasn't used your data, that risk needs to be managed — contractually, organisationally — not waved away as "unlikely."
- The same logic runs in reverse. Anyone training or fine-tuning their own models on third-party data (customers, partners, employees) is exposed to the same kind of question now embarrassing OpenAI — and neither the AI Act nor GDPR distinguishes between "I did it on purpose" and "I couldn't rule it out."
The right question
It isn't "did OpenAI steal two mathematicians' proof." It's: how many active AI contracts in your organisation, right now, would fail to let you rule out with certainty that your data is contributing to the training of a model your competitor will also use?
If answering that means going back to reread a DPA, the work starts there — before the discovery arrives the way it arrived for Buckmaster: after the fact, and in public.
Sources
- Jeremy Kahn, "OpenAI says it cracked Navier-Stokes, one of math's grand challenges", Fortune, September 8, 2026.
- "OpenAI solves longstanding math problem with 10,000-agent swarm — but can't rule out benefitting from a researcher's private Codex data", VentureBeat.
- OpenAI, "Ten advances in mathematics and theoretical computer science".
- "OpenAI's millennium proof dispute raises the question of whether researchers can trust AI labs", The Decoder.
- "NYU Professor Challenges Origin Of OpenAI's AI Proof", Dataconomy, September 9, 2026.
- "A Mathematician Working On The Navier-Stokes Millennium Prize Problem Now Wonders If OpenAI Stole His Notes That He Stored In Codex", WCCFTech.
- "OpenAI ha risolto uno dei problemi matematici del millennio. Polemica nella comunità scientifica", Open, September 9, 2026.
Do you actually know what your AI contracts say?
We help organisations verify the training, data residency and non-disclosure clauses of their AI vendors, and build a risk-management framework for sensitive or proprietary data handed to AI tools.
Talk to us →