Technology

OpenAI Claims Hundreds of Math Breakthroughs. Mathematicians Want Receipts

Martin HollowayPublished 28m ago4 min readBased on 12 sources
Reading level
OpenAI Claims Hundreds of Math Breakthroughs. Mathematicians Want Receipts
source:openai.com

OpenAI released hundreds of claimed solutions to long-standing open problems in mathematics in the week of Oct. 8, 2026. The company said it was evaluating its proprietary models, closed systems owned and run by the company, against open research questions. The response was immediate. Mathematicians at the University of Cambridge and King's College London released a paper that week questioning how frontier labs approach math claims, and Terence Tao criticized OpenAI's approach on social media after the release. TechCrunch

The scale is large and the documentation is partial. OpenAI reported findings on 377 math problems on Oct. 6, 2026. An Oct. 8 account put the manuscript count at 719. Only 10 included the model's chain of thought, the step-by-step record of how the model reached its answer. OpenAI said it had consulted an advisory group of elite mathematicians before the release. It also said it is sharing formalizations of many proofs in Lean, a programming language that lets a computer check each logical step of a proof. TechCrunch

That advisory structure now sits at the center of the dispute. The Advisory Group on Mathematics and Artificial Intelligence is hosted by Princeton University's Institute for Advanced Studies and includes nine prominent researchers. It released guidelines for frontier labs solving math problems at the end of September 2026. The group will advise on review and communication of emerging results and help OpenAI assess their significance. OpenAI

The Cambridge and King's College paper focuses on checkability, whether an outside expert can verify a claim. It documents at least two discrepancies between the natural-language proof, the version humans read, and the Lean code for OpenAI's offered solution to a Navier-Stokes-derived problem. That distinction matters to working mathematicians. A Lean file that typechecks, meaning it passes the computer's logic check, does not settle the readable claim if the two versions do not match.

The current release follows two earlier high-profile claims. OpenAI stated that its solution to the 80-year-old unit distance problem disproved a major conjecture in discrete geometry. OpenAI stated that proof brings in ideas from algebraic number theory. OpenAI In September, OpenAI published an AI-generated solution to the Navier-Stokes Millennium Prize Problem that includes a writeup and a formal proof in Lean. OpenAI A $1 million bounty exists for the first solution to the Navier-Stokes existence and smoothness problem. The BBC reported that OpenAI's claim to have solved parts of the Navier-Stokes equations stirred controversy. BBC

OpenAI has been building toward research-grade evaluation for months. It published First Proof submissions sharing its AI model's proof attempts for a math challenge testing research-grade reasoning on expert-level problems. In August, OpenAI said its new results address long-standing open problems in mathematics and theoretical computer science, including advances in geometry. The friction around that program is not new. Twenty-five leading mathematicians signed an open letter arguing that AI labs are threatening their intellectual work. OpenAI acknowledged its role in an incident where AI agents took over a German wiki forum.

The advisory group's own recommendations set a high bar. Its first request was to stop testing advanced mathematical problems on proprietary models. It suggested OpenAI should help fund the work of human mathematicians needed to make its solutions meaningful. It said it is ultimately up to the mathematical community to assess the extent to which its recommendations were followed successfully.

The broader context here is a mismatch between generation and validation. Labs can now produce proof-shaped papers at a rate that dwarfs normal peer review, much as code assistants can flood reviewers with plausible code. Lean changes the mechanics of checking, but it does not remove the need for readable proofs, released traces, and independent review. Only 10 chains of thought for 719 manuscripts leaves reviewers with outputs and little insight into derivation.

In my view, worth flagging is the specific request to stop using proprietary models on open problems. That request is about contamination and reproducibility. If an open problem is used as a test inside a closed system, the community cannot rerun the experiment, inspect the search, or rule out leakage of the answer into training. Funding for human follow-up is the other half of the same point. A formalization that compiles is a starting point. Turning it into accepted mathematics still requires human labor to align definitions, fill gaps between informal and formal text, and connect the result to existing literature.

For enterprise and research users watching this closely, the near-term lesson looks familiar from code generation. Early output looks plausible, volume rises quickly, and value concentrates where checking can be automated. Lean provides that automation when the formal and informal versions agree. When they diverge, as documented in the Navier-Stokes-derived case, automation does not help. The payoff in the next stretch is less likely to be autonomous prize-winning and more likely to be faster formalization, better proof search, and tooling that makes human mathematicians more productive.