OpenAI Publishes 722 AI-Written Math Papers for Expert Review

OpenAI has published 722 manuscripts grouped into 372 result families, with claimed solutions to long-standing mathematics problems from an unreleased frontier model. The set was listed under the Research post ‘Sharing AI progress in mathematics’ dated October 6, 2026, with the papers hosted in a GitHub repository. The Verge
The repository includes protocols for revising papers and citing them. Each result family groups related papers, so a claim can be updated manuscript by manuscript while the citation target stays stable. Think of it as version control for proofs, the same way software tracks changes file by file.
OpenAI says the average result used the equivalent of three hours of ChatGPT Pro thinking. That puts research output in units of inference time, meaning the computing time a model spends generating an answer. Enterprise teams already track that unit for cost and capacity planning.
The October release expands a September statement in which OpenAI said the same research program had resolved more than 100 long-standing open problems across most areas of mathematics. The later figures replace the earlier count. Output is now presented as manuscripts and families that can be inspected individually.
An August update had shared new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry. OpenAI
One example in the sequence is Navier-Stokes. OpenAI completed the work before human mathematicians, with 10,000 AI systems completing the work in 88 hours. NPR The Guardian
Mathematicians say the Navier-Stokes solution is not telling them much yet. NPR The answer exists. Usable insight lags behind.
The broader context here will be familiar to anyone who has watched automation move into expert work. Generation scales faster than validation. A frontier system can emit hundreds of manuscripts in a single release, but each still requires line-by-line review by people with deep domain training, and peer review was already capacity constrained before machine-generated submissions arrived.
In my view, the revision and citation protocols are the most technically interesting part of the release. They acknowledge that error is expected and that findings will change after publication. That is normal science, but it creates operational questions around tracking versions, keeping a clear record of provenance, and handling downstream use of proofs that may be revised after other work builds on them.
Looking at what this means for research practice, the near-term effect is unlikely to be replacement of mathematicians. It is more likely to be a shift in labor toward checking, refactoring, and explaining machine-produced arguments. Tools for comparing versions of proofs, tracing supporting arguments across families, and isolating the novel step in a long derivation will matter as much as raw solving power. I have seen my own kids learn more from worked examples than from answer keys. If explanation tooling catches up to generation, a library of 372 families could become teaching material and scaffolding for further work.


