Work & Society · Oct 9, 2026
OpenAI withdraws three math manuscripts a day after release as a single sign error topples claims close to the Hodge conjecture
A day after releasing 722 manuscripts at once, OpenAI pulled three of them itself. Human verification cannot keep up with the volume of results produced in a single day, and the claim that "many are formalized in Lean" should be read with caution
Koji Yamamoto · Economics Analyst

Key points
- OpenAI recorded the withdrawal of three manuscripts, a day after release, in the change history (history.md) of its GitHub repository. A single sign error caused claims close to the Hodge conjecture to collapse in a chain
- A proof verified by Lean cannot pass a false statement. So the key parts of the three withdrawn manuscripts were either outside the formalization, or the formalized statements did not match what the manuscripts claimed
- All 722 manuscripts came out at once, and OpenAI found the error itself before outside mathematicians could finish reading them. The remaining several hundred should be treated as "claims" until they are verified
OpenAI withdrew three of the mathematical manuscripts it released on October 6, a day after publishing them. The withdrawal is recorded in history.md, the change history of the GitHub repository (openai/math) where the manuscripts were published (primary source). The cause was a single sign error. The error spread to other manuscripts, including one containing claims close to the Hodge conjecture, and all three collapsed together.
The release comprised 722 manuscripts (372 groups of results), published all at once as the output of having models work on about 4,000 problems. Three manuscripts is a tiny share of that. Still, the withdrawal matters a great deal. The error was found by the publisher before outside mathematicians had finished reading the material, and the release had been promoted on the basis that many of the proofs were formalized in Lean.
One sign error, a chain of collapsing claims
Mathematical results build on other results. If one manuscript gets a single sign wrong, other manuscripts that rely on that lemma lose their footing too. In this case, the chain reached a manuscript containing claims close to the Hodge conjecture. The Hodge conjecture is one of the Clay Mathematics Institute's Millennium Prize Problems.
When human researchers write papers one at a time, the author keeps track of these dependencies, and referees can follow them as a single line of argument. Here, several hundred manuscripts appeared at once, and it is hard for outsiders to trace which manuscript relies on which. Only OpenAI, which has the full set, is in a position to know first how many manuscripts a single error affects. It found out the day after release. The sheer volume itself is what makes verification hard.
How to read "many are formalized in Lean"
When it released the manuscripts, OpenAI said many of the proofs had been formalized in Lean. If a Lean-formalized proof checks, the machine should reject an inconsistency such as a sign error. Even so, the withdrawal happened.
That leaves two possibilities (this is our inference). One is that the key parts of the three withdrawn manuscripts were never formalized. The other is that the formalized statements did not match the claims written in the manuscripts. Either way, the word "many" does not tell readers which results the machine checked, or how much of each. What readers need is a breakdown manuscript by manuscript: how far the formalization goes, and who checked that the Lean statements match the manuscripts' claims. Until that is provided, formalization cannot be treated as a guarantee of correctness for any of the 722 manuscripts.
Formal verification has been promoted as the trump card for handling volumes that human checking cannot keep up with. This episode shows that the trump card means little unless the publisher shows, manuscript by manuscript, how far it actually applies.
Verification cannot keep pace with output
According to reports, each result used computing comparable to ChatGPT Pro, about three hours. By contrast, experts typically need weeks to months to read and check a new proof. Scientific American reported that Andrew Sutherland said "a single agent's claims should be treated as unverified until the model is released" (secondary source). Terence Tao reportedly described the pace as "insane" (secondary source).
The mathematical community has already begun building ways to handle the flood. Tao's blog published "AHM statement on OpenAI's October 6 release of mathematical documents" on October 7 and "What should we tell our students?" on October 8. On LessWrong, a post asks "Are OpenAI's math results creative in an important way?" Before this, the Erdős problems site had narrowed the channel through which it accepts AI-generated proofs that come without explanation. Hexagon, which collects results produced with LLM assistance, also limits submissions to one per day at first. Each of these is an attempt to close, through process, the gap between how much can be produced in a day and how much can be checked in a day.
OpenAI's release reportedly followed advice from an advisory group on mathematics and AI (AGMAI) hosted by the IAS. But AGMAI's role is to advise on how a release is carried out, not to verify individual proofs. Releasing 722 manuscripts at once shifts the burden of verification onto outside mathematicians. This withdrawal makes clear that the burden is real.
How to read the remaining manuscripts
Major claims remain, including on the four-dimensional Kakeya conjecture and progress related to the Riemann hypothesis (both as reported). The withdrawal does not mean the rest are wrong. Finding the error, pulling the manuscripts itself and recording it in the history was, if anything, an honest response.
Even so, the way these results are read needs to change. If three manuscripts fell apart within a day, the remaining several hundred should be treated as OpenAI's "claims" until outside mathematicians have checked them. On formalization, judgment should wait until a breakdown is published showing which theorems in which manuscripts Lean checked, and against which statements. OpenAI also has not yet said which model produced these results. Volume cannot be read as evidence of correctness.
Editorial cartoon

Sources
- https://github.com/openai/math/blob/main/history.md
- https://terrytao.wordpress.com/2026/10/07/ahm-statement-on-openais-october-6-release-of-mathematical-documents/
- https://terrytao.wordpress.com/2026/10/08/what-should-we-tell-our-students/
- https://www.lesswrong.com/posts/neHiCSdL4tJDg7Hrm/are-openai-s-math-results-creative-in-an-important-way