Work & Society · Oct 7, 2026
Erdős problems site drops 'solved' labels and solve counts as unexplained AI proofs pour in, shifting weight to human-readable write-ups and Lean proofs
The standard for what counts as an achievement is beginning to shift from correct answers to understanding that humans can read
Koji Yamamoto · Economics Analyst

Key points
- The Erdős problems website will stop showing 'solved' labels and tallying solved problems, and will no longer record credit for who solved them. The change was announced on Terence Tao's blog on October 6 (primary source)
- The reason is a flood of AI-generated proofs with no explanation. Checking correctness will be left to formalization in Lean, while value is re-centered on human-readable explanations
- Since OpenAI's Navier-Stokes proof was described as 'correct, but humanity has learned nothing,' mathematicians have kept calling for understandable lines of reasoning over headline-grabbing answers. This change is the first case of that view being built into how a site is run
A website that has tracked the open problems left by Paul Erdős, assigning each a number and showing its status, has made a major change to how it is run. It will stop marking individual problems as 'solved' and stop tallying the number of solved problems. It will also stop recording credit for who solved them. Instead, it will prioritize human-readable explanations and formal proofs verified in the proof assistant Lean. The change was posted on Terence Tao's blog on October 6 under the title 'Changes to the Erdős problems web site' (primary source).
The trigger was a surge in AI-generated proofs. The proofs arrive, but they do not explain why they work, and the site was swamped by such submissions. In an era when AI mass-produces answers, a system that attaches a 'solved' mark, lists the solver's name and counts up solutions ends up stamping 'done' on problems that nobody actually understands. The site has removed that stamp altogether.
'Correct answers' are no longer scarce
The Erdős problems site has also served as an informal yardstick for measuring AI's mathematical ability. It has a large number of problems spanning a wide range of difficulty, and each one is marked as either solved or open. For AI companies, it was ideal material for claiming they had solved a certain number of problems. Dropping the solve count means that yardstick can no longer be used.
The backdrop is a series of events over the past few weeks. In September, OpenAI announced that an internal model had solved a problem concerning the Navier-Stokes equations and had dispatched more than 100 other long-standing open problems. It also released a 166-page manuscript and a Lean formalization that compiles, and the prevailing view is that the work is technically correct. Even so, the Clay Mathematics Institute has kept the problem on its list of unsolved problems.
NPR reported on September 22 under the headline 'AI solved one of math's hardest problems. Humans have learned nothing (for now).' James Maynard of Oxford said that 'extracting human understanding from this new AI proof has so far been very difficult.' Javier Gómez-Serrano of Brown University said 'this paper was not written for humans,' while Tristan Buckmaster of NYU called it 'a terribly written paper.' Correctness is no longer in dispute. The question is what humans have gained from that correctness.
Leave correctness to Lean, ask humans for understanding
What stands out about the change is that it splits the work of evaluation in two. Whether a proof is correct can be checked mechanically through formal verification in Lean. That removes the burden of mathematicians reviewing AI proofs line by line. What the result means as mathematics, however, cannot be conveyed without a human-readable explanation. The site's decision to shift its emphasis from answers to explanations reflects a judgment that this second element is what the mathematical community truly wants to build up.
This idea has been argued repeatedly in a series of guest posts on Tao's blog since September. Tapio Schneider of Caltech wrote in 'Headlines and inside stories' on September 24 that a proof that is correct but cannot be understood provides only a 'headline' and not the 'inside story.' Amit Sahai argued that more mathematicians should be trained so that humans can verify AI results. Antonio Auffinger of Northwestern warned that the acceleration of proofs by AI is eroding mathematics' culture of sharing. On October 5, the day before the announcement, the same blog published 'The future of mathematics.' The change to the Erdős problems site can be seen as the first concrete example of these debates making their way into how a site is actually run.
'Who solved it' loses its meaning
Ending the record of credit is also significant. In mathematics, the name of the first person to solve a problem has long been the unit of achievement. But in a workflow where an AI produces a draft, a human polishes it and another AI formalizes it, even answering 'who solved it' becomes difficult. A guest post by Diego Córdoba and Luis Martínez-Zoroa on the same blog on October 4 explicitly cited recent work on the Euler and Navier-Stokes equations that used large language models. AI-produced results are already among the prior work being cited.
At this point, counting named 'solutions' no longer works well either as a yardstick of how far mathematics has advanced or as a system for recognizing individual achievement. The site has stepped away from both. What it will keep is a record of what humans have understood about each problem.
Mathematics begins to redefine 'achievement'
The change goes beyond one site's operating policy. It shows the mathematical community beginning to redefine, on its own terms, what it recognizes as an achievement. AI companies have used the number of problems solved to advertise their models' performance. An open letter issued on September 11 by 25 Fields Medal winners criticized the use of open problems to promote AI benchmarks. Mathematicians have now responded by changing what gets counted in the first place.
Questions remain. Who will write the Lean proofs and the human-oriented explanations, and how many? Once AI starts writing the explanations too, should those count as 'human-readable understanding'? The practical details of how to record and evaluate human understanding have yet to be tested. Still, the direction is clear. In a world where answers have become cheap, the mathematical community has begun choosing to build up understanding rather than pile up answers.
Editorial cartoon
