AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Jul 16, 2026, 04:01 JST3 min read

GPT-5.6 Sol disproves 30-year statistics conjecture in 90 minutes

GPT-5.6 Sol disproves 30-year statistics conjecture in 90 minutes

Key takeaway

  • A University of Pennsylvania professor used OpenAI's GPT-5.6 Sol Pro to disprove a 30-year-old statistical assumption in 90 minutes—something mathematicians had failed to do.

  • The Benjamini-Hochberg procedure, a method cited over 130,000 times, was assumed to work reliably with correlated data, but the AI showed cases where it misses its target false discovery rate.

  • Although the practical gap is small, the result highlights AI's growing problem-solving capability in mathematics and raises questions about whether AI can reason to genuinely new knowledge.

3 Key Points

  1. What happened

    A University of Pennsylvania statistics professor used OpenAI's GPT-5.6 Sol Pro to disprove a longstanding assumption about the Benjamini-Hochberg procedure, a widely-used statistical method for controlling false positives. The AI constructed a statistical model showing the method's false discovery rate can exceed its target level when data is correlated and normally distributed—something mathematicians had assumed but never proven in 30 years. GPT-5.6 Sol Pro completed the work in about 90 minutes; GPT-5.5 could not find a solution even after roughly 20 hours.

  2. Why it matters

    The Benjamini-Hochberg procedure, introduced in 1995, has been cited over 130,000 times and is used across modern statistics and scientific fields to filter out false alarms when testing thousands of hypotheses at once. While the gap the AI found is small in practical terms (0.104 versus the target 0.1), Berkeley statistician Will Fithian called the disproved conjecture 'the most interesting open problem in my area of statistics.' The result signals AI's advancing capability to solve problems that eluded human experts, marking what Fithian described as 'another marker of advancing AI capabilities whose consequences will reach far beyond math.'

  3. What to watch

    The AI's solution combined known statistical methods in an unusual way rather than inventing entirely new ones, leaving open the question of whether AI systems can generate genuinely novel knowledge or only recombine what they learned during training. Dobriban notes the result 'mainly matters for theory at this point' and that practical effects need further study. The full chat and code are publicly available.

Ask the AI about this article →

Context & Analysis

The Benjamini-Hochberg procedure has been a cornerstone of modern statistics since 1995, particularly in fields like genomics where researchers must filter false positives from thousands of simultaneous tests. For nearly three decades, the statistical community assumed the method would work reliably even when data points are correlated—a common real-world scenario, such as when genetic variants are inherited together. However, no one had formally proved this assumption. The University of Pennsylvania's Edgar Dobriban leveraged GPT-5.6 Sol Pro to fill that gap, constructing a counterexample showing that under certain conditions, the false discovery rate can exceed the target level. The speed of the solution is what stands out: 90 minutes versus roughly 20 hours for the previous generation model, and decades of unsuccessful human effort. Dobriban acknowledged that the AI's approach combined existing statistical methods in an unusual way rather than inventing new ones—a pattern seen in similar AI breakthroughs in mathematics. This raises a fundamental question about AI's reasoning: can these systems generate truly novel knowledge, or are they limited to sophisticated recombination of ideas learned during training? Even if recombination is the ceiling, Dobriban's work demonstrates practical value as a tool embedded in human workflows. The broader implications remain uncertain, though some researchers, including deep learning pioneer Richard Sutton, believe genuine self-improvement and generalization require capabilities beyond recombination.

FAQ

What is the Benjamini-Hochberg procedure?
It is a statistical method developed in 1995 to control false discovery rate—the share of reported significant results that are actually false alarms—when researchers test thousands of hypotheses at once, such as scanning the human genome for disease-linked genes. The original paper has been cited over 130,000 times.
How long did GPT-5.6 Sol Pro take to solve this problem?
About 90 minutes. In contrast, GPT-5.5 could not find a solution even after roughly 20 hours of work with several agents.
Does this mean the Benjamini-Hochberg procedure is unusable?
No. The gap the AI found is relatively small (0.104 versus the target 0.1), so the result mainly matters for theory at this point. Practical effects still need further study, and the finding does not mean the procedure is generally unusable.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 2h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 2h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCohere VP: Enterprise AI sovereignty needs full control of agent stack