OpenAI
OpenAI Releases 722 AI-Written Math Papers Claiming Hundreds of Breakthroughs, but Experts Urge Caution

SAN FRANCISCO — OpenAI has released 722 mathematical manuscripts produced by an unreleased artificial intelligence model, claiming solutions to hundreds of long-standing open problems in a move that has both impressed and unsettled the mathematics community.

The manuscripts, published Tuesday in a public GitHub repository, cover 372 "result families" that group related papers together, according to the company. The release is the largest public display yet of an AI system's ability to tackle research-level mathematics, but many of the results have not been independently verified.

The Advisory Group on Mathematics and Artificial Intelligence, an independent panel of mathematicians known as AGMAI that OpenAI consulted on the release, said the batch includes solutions to "hundreds" of open questions. The group has not endorsed the findings, according to reports.

How the results were produced

OpenAI said the unreleased frontier model was posed roughly 4,000 problems during an internal evaluation. The 722 manuscripts represent the results the company judged significant enough to publish.

The company said the "average result" used computing power equivalent to about three hours of ChatGPT Pro thinking time.

Alongside the papers, OpenAI released summaries of some of the model's reasoning, estimates of the computing resources used and statistics on the number of problems attempted. According to Scientific American, the company did not publish the prompts used to generate the results. OpenAI has also not disclosed the name of the model.

"For this release, we're publishing the results in a GitHub repository, with protocols for paper revisions and citations," OpenAI said. The manuscripts were released under an open-source license.

Claimed results on famous problems

According to reports summarizing the release, the manuscripts include claimed results on several well-known problems, including a proof of the Unique Games Conjecture in theoretical computer science and a resolution of the free group factor problem, a question in operator algebras that has been open since the 1940s.

The collection also reportedly includes a claimed result showing that the Riemann zeta function has no zeros in a region where the real part exceeds 7/8, a weaker version of the famous Riemann Hypothesis that some have described as a "quasi-Riemann hypothesis."

These claims have not been confirmed by the mathematical community. Major results in mathematics typically require months or even years of careful review by experts before they are accepted. OpenAI has acknowledged that the level of verification varies across the results, and that some proofs that have not been formally checked may contain errors. Some proofs have been formalized in Lean, a computer program used to verify mathematical arguments.

A run of breakthroughs

The release had been expected for weeks. In September, OpenAI said its internal model had "resolved more than 100 long-standing open problems across most areas of mathematics," but did not say which problems it had solved or when the results would be published.

That announcement drew criticism from some mathematicians, who complained about the lack of detail and the difficulty of verifying claims made without published proofs.

The latest release extends a series of AI advances in mathematics. In recent years, AI systems from OpenAI and Google DeepMind have achieved gold-medal-level performance at the International Mathematical Olympiad, a competition for high school students. The new results move the focus from competition problems, which have known answers, to open research questions, where verification is harder and slower.

Advisory group's guidelines

AGMAI was formed in September and is hosted by the Institute for Advanced Study in Princeton, New Jersey, according to reports. Its members reportedly include Fields Medal winners Timothy Gowers, Martin Hairer and physicist Edward Witten.

The group published its first recommendations in late September, urging AI labs to release mathematical results promptly and through established academic channels where possible. It called on companies to disclose details such as the name of the model used, the prompts and the computing costs.

The group also urged AI companies to "refrain from treating the release of mathematical results as marketing vehicles to promote their models," a practice it said causes significant harm to the mathematical community.

OpenAI said it developed its release process in consultation with the advisory group and the Institute for Advanced Study. The company is also considering community-hosted options for the results that would comply with the group's recommendations.

Questions about ethics and credit

The release has raised broader questions about research ethics, academic credit and the future of mathematical research.

Mathematicians have raised concerns about how to assign credit for AI-generated results, how to verify such a large volume of work and whether a flood of AI-produced papers could overwhelm the peer review system.

There have also been disputes over OpenAI's communication with mathematicians. Some attendees at earlier meetings recalled that the company had indicated it would not release all of its results at once. OpenAI spokesperson Lindsay McCallum said the company was "not aware" of any such promise, according to reports.

Because the model has not been publicly released, other researchers cannot independently reproduce its reasoning process, a limitation critics say makes it harder to evaluate the results.

Competition in AI

The release comes amid intense competition among AI companies to demonstrate advanced reasoning capabilities. Mathematics has become a key testing ground because results can, in principle, be checked rigorously.

OpenAI has said it plans to continue testing its most advanced models on mathematics and other scientific problems, arguing that AI could give researchers powerful new tools for discovery.

Mathematicians are expected to begin the lengthy process of reviewing the manuscripts. Results that withstand scrutiny could reshape several fields, while any errors could fuel skepticism about AI-generated research.

The full impact of the release may take months or years to assess, as experts work through hundreds of dense papers spanning many areas of mathematics.