OpenAI Releases Nearly 400 AI-Generated Math Results for Review

OpenAI released nearly 400 new mathematical results produced by an internal AI model, spread across 719 manuscripts.
The collection was described in an OpenAI Research post titled "Sharing AI progress in mathematics" dated October 6, 2026, and shared through a large GitHub repository. OpenAI lists the post in its newsroom and says the work consists of new results that address open problems in mathematics. OpenAI
The scope is broad. The manuscripts cover combinatorics, geometry, number theory, theoretical computer science, algebra, topology, probability and statistical mechanics, and mathematical physics. OpenAI published guidance for navigating the repository, which sorts outputs by field and result type. The Verge
Verification status varies across the set. OpenAI said the results are at different stages of verification and that many, but not all, manuscripts have been formalized, meaning translated into a strict form a computer can check. Of the 719 manuscripts, the company said 300 top-line results had been formalized, about 42 percent. It said it will update the repository with more formalizations as it obtains them.
Some manuscripts include formalizations in Lean. Lean is proof-checking software that checks each logical step apart from the informal written explanation. The repository keeps the informal writeups separate from the machine-checkable files where those files exist.
How the results were generated is part of the disclosure. OpenAI tests its models on open research problems as part of model development. After performance on its existing math tests leveled off, the company expanded testing to open research problems. Some outputs build upon earlier results produced by its models, which points to chained work rather than only isolated attempts. OpenAI
Early coverage gave slightly different totals. The New York Times reported findings on 377 math problems, while The Economist and Scientific American cited 372 mathematical results. The latest accounting puts the release at nearly 400 results in 719 manuscripts, with 300 top-line results formalized. The difference reflects manuscript count versus distinct results, plus updates to the repository after October 6. The New York Times
The Economist also reported that OpenAI's AI agents solved long-standing mathematical problems in weeks.
Outside review has started. The Verge spoke with more than three dozen mathematicians about the release. Two of those contacted were Alvaro Lozano-Robledo, a professor of mathematics at the University of Connecticut, and Kevin Buzzard, a mathematics professor at Imperial College London working in algebraic number theory.
The release follows earlier mathematics disclosures from OpenAI. On September 8, 2026, the company shared an AI-generated solution to the Navier-Stokes Millennium Prize Problem including a writeup and a formal proof in Lean. On May 20, 2026, it said a model solved the 80-year-old unit distance problem, disproving a major conjecture in discrete geometry. On September 21, 2026, OpenAI created an Advisory Group on Mathematics and Artificial Intelligence to advise on review and communication of emerging results and assess their significance.
The broader context here is testing and review. When standard math tests no longer separate strong models, a lab can turn to unsolved research problems as its test. That changes the work for outside reviewers. Mathematics already depends on careful line-by-line reading. A release of hundreds of claimed results, with less than half in machine-checkable form at publication, puts the first burden on reading informal drafts, with computer checking to follow. The Economist account of problems solved in weeks will get close attention from researchers used to measuring progress in months and years.
In my view, the figure to watch is that 42 percent formalized, or 300 of the top-line results. Formalization alone does not settle correctness, because the formal definitions can still differ from the intended theorem, but it shrinks the places where mistakes can hide and allows automatic rechecking. For research and engineering teams, the test will be whether the Lean files plus drafts shorten the time to independent reproduction. Without that, scale adds review work.
What matters for model development here is the chaining detail, where some outputs build on earlier model-produced results. In this step-by-step method, small results from one run become starting points for the next, which tests whether a system can build a library over time rather than solve single problems in isolation. If reproducible, it fits with agents that keep mathematical context across weeks of work rather than single answers. Worth flagging, human mathematicians still have to judge novelty, relevance, and whether the formal claim matches the open problem as stated.
Looking ahead, the optimistic case is clear. Mathematics that can be checked by machine, shared openly with code and explanations, gives the field more to test, reuse, and correct. Even partial formalization builds a public set of examples for improving proof tools, search over math libraries, and automated tactics. Over time, that could lower the cost of checking hard proofs and leave mathematicians more time for choosing problems and conceptual work.


