722 "Manuscripts" Released on GitHub
On October 6th, OpenAI released a research paper titled "Sharing AI progress in mathematics," along with a collection of mathematical results generated by its previously unreleased in-house frontier model. The files are located in the "openai/math" repository on GitHub under the Apache-2.0 license. The collection consists of 722 manuscripts, organized into 372 related "families." The topics covered are broad, including number theory, computational complexity, geometry, and mathematical physics. While the magnitude of the mathematical achievements is significant, what caught my attention as a researcher was the "method" of presenting the results. How do you gain the trust of the mathematical community for AI-generated mathematics? This collection reveals OpenAI's exploration of this process, often with external advice.
Advice from the Institute for Advanced Study Group
OpenAI explains that it consulted with an independent "Advisory Group on Mathematics and Artificial Intelligence" at the Institute for Advanced Study in Princeton regarding this publication method, and referred to their previously published recommendations. There is already a movement from mathematicians to standardize the publication practices for AI-generated results, and OpenAI designed its publication method in line with this. Specifically, it has established protocols for manuscript revision and citation, and ensures that older versions remain accessible. Regarding the paper's hosting, in addition to GitHub, they are exploring community-based hosting options that align with the advisory group's guidelines.
The Meaning and Limitations of "Lean Verified"
The Lean-based formalization is prominently featured as a guarantee of reliability. Lean is a programming language that allows mathematical proofs to be mechanically checked by computers. A formalized proof will fail from the start if there are leaps or holes in the logical flow. OpenAI has published Lean versions of many proofs and plans to add more as they become available. However, it's important to note that this is "many," not "all." Furthermore, as a general principle, Lean guarantees that "formalized propositions have been proven," but whether those propositions precisely match the theorems intended by humans needs to be verified separately by humans. Mistakes at the formalization stage cannot be detected by machine inspection alone.
Seeing the "Contents" of Transparency in Numbers
OpenAI also provides information on the process leading to the results. This includes 10 documents summarizing the model's inference, an estimate of the computational cost converted to the use of ChatGPT Pro, and statistics on the number of problems attempted. According to the report, approximately 4,000 problems were involved in the evaluation, and the average computational cost per result is equivalent to about 3 hours of thinking on ChatGPT Pro. The README also states that the evaluation was extended to unsolved research problems because existing mathematical evaluations had reached saturation. What we need to calmly consider here is that we cannot simply divide approximately 4,000 problems by 372 families to obtain a "success rate." This is because one problem may lead to multiple documents, and vice versa. While the disclosure of the denominator of the trials is commendable, caution is needed in interpreting it.
The Biggest Constraint: Inability to Reproduce
On the other hand, the biggest constraint from a verification perspective is clear. The model that produced the results remains undisclosed, making it impossible for external mathematicians to reproduce and verify the same procedure. OpenAI itself only states that it is working to release this model in a responsible manner. In other words, at present, anyone can verify the parts that have been machine-checked with Lean, but it is impossible to verify from the outside "how the proof was arrived at" or "whether similar results can be obtained for other problems under the same conditions."
Self-Identified Manuscript Quality as a Challenge
Another point that reveals their frankness is the list of areas for future improvement. OpenAI has promised to improve the quality of citations, the way mathematical explanations are presented, and the way results are presented, in preparation for future publication. Conversely, this means that there is room for improvement in these areas of the current manuscript. A research paper is not sufficient if the proof is correct; it can only be used as community knowledge when it includes context within previous research and explanations that readers can understand. Given the sheer volume of 722 papers, whether the quality aligns with the volume remains to be seen, and we must await individual evaluations by mathematicians.
What Researchers Should Keep an Eye On
Ultimately, the value of the results will be determined by how mathematicians in each field evaluate each paper. OpenAI has also indicated plans to fund workshops, conferences, and special programs focused on understanding the major achievements of AI. What's important about this release is not the sheer size of the results themselves, but the fact that a "framework" has begun to emerge for presenting AI-generated mathematics in a way that the community can verify, cite, and critique. We will continue to closely monitor whether the Lean formatting is expanded, models are made available externally, and verification results from independent mathematicians become readily available.