AIToday
Large Language ModelsLatent SpacePublished: Oct 7, 2026, 16:00 JST

OpenAI posts 722 math papers, incl. Quasi-Riemann Hypothesis

OpenAI posts 722 math papers, incl. Quasi-Riemann Hypothesis

3 Key Points

  1. What happened

    OpenAI published 722 math manuscripts in 372 result families to a public GitHub repo, drawn from about 4,000 problems and averaging roughly 3 hours of ChatGPT Pro thinking compute each. Levent Alpöge called it "the most significant moment in mathematical history".

  2. Why it matters

    If the results hold, a model OpenAI has not released may have solved many of the top 500 open math problems at a compute cost one observer called "not much" — a possible shift in who can produce frontier mathematics.

  3. What to watch

    Independent verification, since one analysis estimates about 20% of the results are disproofs or counterexamples, and Will Depue expects some to fail scrutiny. He built citedbyagi.com to track which human papers the release cites.

WHO IT HITSMathematics and research-adjacent teams that rely on peer review should expect a backlog of AI-generated results to check. AI lab leaders weighing release policy face the question of what to do with a capable model whose outputs are published but whose weights stay private.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The release is framed as a broad set of mathematical results from an internal frontier model, not a product launch. OpenAI says it consulted the Institute for Advanced Study's independent Advisory Group on Mathematics and AI on how to release them — a step that points to the awkwardness of publishing findings from a model that stays private.

The scale is unusual: 722 manuscripts grouped into 372 families, drawn from an evaluation of about 4,000 problems. Among the highlighted items are a result for integer multiplication faster than n log n and a uniqueness result for the elastic inverse problem that the paper says had been open in 3D since 1994. Commenters also point to partial progress on Riemann, Hodge and BSD.

What the outcome hinges on is scrutiny. Will Depue expects some results should not survive it, and an analysis arguing that roughly 20% are disproofs or counterexamples cuts against the idea that these are mostly brute-force wins. Levent Alpöge's praise came alongside his own note of reported scooping and conflict-of-interest problems involving other labs' users — a sign the social machinery around AI math is straining as much as the mathematics.

FAQ
Can I use the model that produced these results?
No. The release includes papers, proof artifacts and selected reasoning summaries, but the model itself remains unreleased.
How much compute did each result take?
Roughly three hours of ChatGPT Pro thinking compute per result on average, from an evaluation of about 4,000 research problems.
Have the results been independently checked?
The notable claimed results were reported by individual commentators and have not been independently verified. One analysis estimates about 20% are disproofs or counterexamples.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle's Gemini Nano Banana 2.1 halves image output to $30