Work & Society · Oct 8, 2026
OpenAI releases 722 math manuscripts without naming its internal model, and mathematicians build their own safeguards
The results keep piling up, but OpenAI has not shown how they were produced. Before mathematicians can check any single result, they are rebuilding the systems that will have to receive them
Koji Yamamoto · Economics Analyst

Key points
- OpenAI has released 722 mathematics manuscripts produced by an internal model, without naming the model
- Mathematicians have criticized OpenAI for not disclosing how the manuscripts were generated and selected, or how they were verified
- Mathematicians are already building systems to receive such results, including rule changes at the Erdős problems site, the AGMAI advisory group and a statement from ICIAM
OpenAI has released 722 mathematics manuscripts produced by an internal model (OpenAI blog, primary source). The company did not say which model it used. It also did not describe how the manuscripts were generated, how it chose which results to release or how far they had been verified. That gap is the main focus of criticism from mathematicians. Before anyone can debate whether the results are correct, it is unclear what was produced and how it came about. Mathematicians have therefore started building their own systems to receive and verify such results before they turn to checking individual ones.
The release shows "how many" but not "how"
When OpenAI announced its Navier-Stokes result in early September, it said its internal model had resolved more than 100 open problems in mathematics. The 722 manuscripts package that work as papers and make it public. Scientific American reported that hundreds more results had been released into a field that was already in shock.
But OpenAI released only the outputs. The manuscripts do not give the model's name, the compute used, the number of discarded attempts or the stages at which humans intervened. For Navier-Stokes, OpenAI itself disclosed the scale of the effort: about 10,000 agents running for 88 hours. This time it has not done even that. In mathematics, how a result was obtained matters alongside whether it is correct, because evaluation and citation depend on both. Without knowing the process, someone reading the manuscripts one by one cannot tell whether a given paper is a carefully selected best case or typical of the whole set.
Reactions have varied widely. The economist Tyler Cowen of George Mason University wrote just a few words on his blog, Marginal Revolution.
OAI doing math again
The word "again" shows how the reception has shifted. Navier-Stokes was met with astonishment, but a batch of several hundred manuscripts is now starting to be treated as routine. Meanwhile, the work of absorbing all of these results has not gone down. If anything, it has grown. Gary Marcus has also published commentary on the release, and the newsletter AINews covered it as well.
The criticism is about "how it was made" more than "whether it is right"
Mathematicians are not claiming that OpenAI's results are wrong. Most experts also considered the Navier-Stokes proof technically correct. Their objection has been about whether the results are presented in a form others can verify. In September, Martin Hairer wrote that the mathematical results released by AI companies fall far short of acceptable mathematical practice. Javier Gómez-Serrano of Brown University said the Navier-Stokes paper was not written for humans.
A batch of 722 manuscripts makes the problem bigger. Experts can take the time to study one major proof closely. With several hundred, they need information just to decide which ones to read first. Which model produced each result, out of how many attempts and on what criteria was it selected? Without that information, mathematicians cannot decide where to spend their verification effort.
Mathematicians are building the "receiving end" first
Mathematicians are therefore changing how results are received before they rule on any individual result. Every development of the past few weeks has moved in that direction.
The Erdős problems site closed its comments and changed how credit is assigned
On October 6, Terence Tao's blog carried a guest post by Thomas Bloom. The Erdős problems site had become a place where AI-generated proofs with no explanation were posted to claim priority. In response, the site stopped accepting new comments on individual problems and now takes submissions by email. It has also stopped counting the number of solved problems and will no longer credit results to whoever solved them, human or AI. Instead, it will give more weight to human-readable explanations and linked Lean proofs. In effect, the site rebuilt the way it receives and records results to cope with the volume of AI output.
AGMAI began with advice on "how to release"
AGMAI, an advisory group on mathematics and AI launched in September, was set up as an independent body after it declined OpenAI's offer to form an external advisory board. Its first task is not to verify results. It is to advise on how OpenAI should coordinate the release of the many results it reports its internal models have produced. The 722 manuscripts fall squarely within that remit. However, AGMAI's charter does not set out how verification should be done. OpenAI has also said that the pace of its internal progress is outside the scope of the group's advice.
Calls for verifiable environments and transparency on compute
In a statement dated September 24, the International Council for Industrial and Applied Mathematics (ICIAM) called for environments where anyone can verify results, transparency about the compute used and fair access. Compute and process are exactly what OpenAI has left undisclosed this time. Amit Sahai has argued that, so that humans can check AI results, the field should build up a large pool of mathematicians as an intellectual reserve that can be mobilized at any time. Harvard has also held a summit on doctoral education in mathematics in the age of AI. Tao's blog has continued to carry posts on AI and mathematics (October 6, October 7).
Who pays for verification?
What these efforts share is a common pattern: because the producers do not disclose their process, the people receiving the results are building systems to fill the gap. Formalization in Lean, the emphasis on human-readable explanations, limits on submission channels and the training of reviewers all add up as costs paid in mathematicians' time and by their institutions.
OpenAI has released 722 manuscripts, but the work of deciding what place these results have in mathematical knowledge has been left to the mathematical community. Until the model's name and the method behind the results are disclosed, mathematicians will keep putting their effort into the systems that receive such results before they can turn to whether the results are true.
Editorial cartoon

Sources
- https://openai.com/index/sharing-ai-progress-in-mathematics
- https://www.scientificamerican.com/article/openai-unleashes-hundreds-more-math-results-upon-a-field-already-in-shock/
- https://www.latent.space/p/ainews-quasi-riemann-hypothesis-openai
- https://garymarcus.substack.com/p/complementary-remarks-from-gary-marcus
- https://terrytao.wordpress.com/2026/10/06/hexagon/
- https://terrytao.wordpress.com/2026/10/07/the-barriers-of-perception/