Physics AI Benchmarks May Be Broken, New Audit Argues

A preprint posted on 11 Sep 2026 argues that leading physics benchmarks understate what frontier models, the most capable current systems, can do on well-posed problems, questions that are clearly stated with one correct answer.
The paper, "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks", is listed on arXiv as 2609.13009 and names Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin and Yu He among its authors. arXiv
How the check was done
The authors state that their expert audits established or corrected reference solutions to provide ground truth for evaluation, in particular for CritPt and CMT-Benchmark. They treated the answer key and the grading software as items to audit, not as fixed infrastructure. arXiv
The preprint defines two types of failure. A grader error is a well-posed problem with a correct reference solution where the model gives a correct answer but the evaluator marks it incorrect. A benchmark error is a defective problem statement or reference solution, including incorrect reference, inconsistent conditions, ambiguity, or missing assumption.
What the new scores show
The preprint states that Anthropic obtained expert corrections to 31 problem statements and evaluated an internal benchmark version called CritPt-Corrected. On that corrected version, Anthropic's Fable 5.1 achieved mean@16 of 88.4%, an average score when the model is allowed 16 attempts per problem.
The same audit approach was applied to other physics suites. The audit excluded 13 benchmark-error questions from its 100-question PHYBench subset, leaving 87 questions. It excluded 26 benchmark-error questions from its 100-question PRISM-Physics subset, leaving 74 questions. It excluded 18 benchmark-error questions from its 100-question UGPhysics subset, leaving 82 questions.
The HLE-Physics slice had a higher exclusion rate. The preprint states its HLE-Physics audit excluded 86 benchmark-error questions and retained 116 questions. Corrected scores were then obtained by running the HLE-adapted evaluation pipeline on the retained subsets.
TPBench appears as related context. The dataset is associated with arXiv ID 2502.15815 and website tpbench.org, and was constructed to benchmark and improve AI models. Stanford Q-FARM
Why this matters
The broader context here is familiar to anyone who has maintained a test. Grading software can fail without warning. Answer keys can drift. Unclear questions can persist because models learn to work around them. A grader error penalizes a capable model, while a benchmark error makes the task itself unclear.
In my view, the practical point is not that physics is solved. Uncorrected leaderboard numbers are a weak guide to ability on clearly specified physics. Mean@16 on a corrected set measures something different from single-sample accuracy on a noisy set. Filtering to well-posed questions raises scores by construction, but it also changes what the score means. The near-saturation language should be read narrowly. It applies to well-posed, gradable problems under expert-corrected references, not to open-ended research physics. Better measurement usually leads to better systems, and physics is a field where correct answers can, with care, be checked.


