Newsletter
OpenAI has published 722 mathematical manuscripts on GitHub, which belong to 372 different families of results, all stemming from an internal model that has not yet been made public.
The spokesperson, OpenAI, stated that almost all of this content comes from a set of prompts provided to a single agent, AI, although some of it may have involved multiple attempts.
This is a bold claim, and it could also represent an important breakthrough in the field of mathematics. However, not everyone is convinced by it.

“Unless they make the model public and people are able to reproduce the results, I think any claim that a single agent can solve a problem in one go should be considered unproven,” said mathematician Andrew Sutherland to Scientific American. “We should demand evidence,” he said.
OpenAI released a concise summary of reasoning for 10 of those results. Meanwhile, according to the formalized catalog in the repository, only 162 out of these 722 papers contain main results that have been checked by a computer, which accounts for approximately 22% of the total content. These results were translated into Lean – a software that mechanically checks each step of logical deduction.
OpenAI himself also stated that not all manuscripts are in the Lean formalized format, and mentioned that 'there may be issues with some of the unformalized results.' In other words, a lot of the content they publish could be incorrect.
Passing the check for Lean only proves that the statements made in Lean are valid, but it does not prove that these statements are completely consistent with the original question, nor does it prove that the result is new or significant. And this is precisely what mathematicians must now judge.
And this is also where researchers start to frown.
“I tried to read the proof by OpenAI regarding planar colors with a count of 6 or more, but it’s completely unbelievable alien mathematics? The model actually found any K-coloring to be equivalent to ‘weakly measurable’ K-coloring, which seems to have appeared out of nowhere.” Dmitry Rybin wrote on X.
“Repository openai / math has been closed, and it has never accepted any pull or request. This is very disappointing. If you post 722 manuscripts and require them to be in the Lean format, then there needs to be a place for people to submit them. I am using Lean to formalize proofs of conjectures about Saxl…” Keith Adler wrote on X.
"The current situation is that AI is able to produce mathematical arguments in situations that humans cannot understand, verify, or be responsible for," stated the Princeton Institute for Advanced Study in a statement. "We believe that human understanding of mathematics remains crucial. In this new era, how can we strive to establish a new paradigm that incorporates human understanding of mathematics into responsible academic outcomes?"
However, there are also those who hold a more optimistic attitude. Professor Abhishek Saha from the University of Toronto wrote, “Today is a very important day for mathematics.” But he pointed out that most of these issues represent “extraordinary progress within existing research plans” or “surprising breakthroughs.”
This means that most of the problems in this set are quite interesting, but they are not as impossible or as significant as solving the Millennium Problems, which could transform our understanding of mathematics. Among the 722 problems, there is only one that falls into that category: the Riemann Hypothesis ( Quasi - Riemann Hypothesis ).
"I still have some further thoughts regarding the 372 results published today by OpenAI, which cover 722 manuscripts. If I were to classify them according to the breakthrough nature of the theorems proven and published by mathematicians, I would roughly divide them into four categories: A) Non-breakthrough..." Abhishek Saha wrote on X.
University of Toronto mathematician Daniel Litt holds the opposite view, believing that there is no reason to require companies to keep the answers to these mathematical problems confidential.
Last month, a different approach was taken to the proof of Fermat's Last Theorem, which had been checked by Lean, and the entire 13 million lines of code were publicly released on GitHub. That proof formalized the theorem published by Andrew Wiles in 1995, rather than claiming to present new results.
OpenAI indicates that as more formalized versions in the form of Lean are obtained, the company will continue to supplement them; currently, among these 722 manuscripts, 162 already have the Lean format.












