Technology

OpenAI Claims AI Progress on Open Math Problems, With Code to Check

Martin HollowayPublished 19m ago3 min readBased on 4 sources
Reading level
OpenAI Claims AI Progress on Open Math Problems, With Code to Check
source:openai.com

OpenAI published new results on open problems in mathematics from an internal frontier model. The announcement, titled "Sharing AI progress in mathematics," is dated October 6, 2026. OpenAI

The publication included Lean proof formalizations for the advances. Lean is software that checks each step of a proof, much as a compiler checks code. Related research details were shared on GitHub. OpenAI Research

People and background

Asaf Karagila identifies himself as the owner of the karagila.org homepage. He states his current posts as University Academic Fellow at the University of Leeds and holder of a UKRI Future Leaders Fellowship. He lists his ORCID iD as 0000-0003-1289-0904. Asaf Karagila

He earned B.Sc. and M.Sc. in mathematics from Ben-Gurion University of the Negev from 2007 to 2012 under advisor Uri Abraham. He earned a Ph.D. in mathematics from the Hebrew University in Jerusalem from 2012 to 2017 under advisor Menachem Magidor. He then served as a postdoctoral project assistant at Technische Universität Wien from 2017 to 2018 supervised by Martin Goldstern. He was a Newton International Fellow at the University of East Anglia from 2018 to 2020 hosted by David Aspero. He was a UKRI Future Leaders Fellow at the University of East Anglia from 2020 to 2022. He has been a University Academic Fellow at the University of Leeds since 2022.

The Partition Principle (PP) is an axiom introduced by Russell. B. Venkataramani authored a 2021 work titled Axiom of Choice and the Partition Principle discussing the Partition Principle in the context of its similarities and differences with the Axiom of Choice. The Axiom of Choice is a standard rule about choosing elements from sets. McMaster Thesis

Checking the claim

The broader context here is how such claims get checked. Informal mathematical text leaves gaps for readers to fill. Lean code does not allow that looseness. It must typecheck, which means the computer confirms every step follows the rules. That property is useful. It shifts scrutiny from prose interpretation to artifact inspection.

Worth flagging is what the combination of formal proofs and code release enables for specialists. Readers with the toolchain, the setup needed to run Lean, can load the definitions and follow each step without trusting the narrative summary. They can test lemmas, small supporting results, in isolation. They can look for gaps between the stated open problem and what was actually formalized. Precision matters here. Open problems often turn on exact hypotheses.

In my view, the interesting question is not whether a model produced text that looks like mathematics. Output fluency is cheap now. The interesting question is whether the formal artifacts hold up under independent checking and whether they address the problems as specialists understand them. That takes time. Set theory, and choice principles in particular, reward careful distinctions. A small change in assumptions can change what a result says.

Looking at what this means for working mathematicians and tool builders, the near-term value lies in workflow. Shared formalizations give experts something concrete to probe, reuse, and correct. GitHub access supports that loop. It allows issue tracking, version comparison, and direct reference to specific lines. The long arc is positive. Better checking infrastructure helps humans and models alike. It lowers the cost of finding errors early.