
A University of Pennsylvania professor used OpenAI's GPT-5.6 Sol Pro to disprove a 30-year-old statistical assumption in 90 minutes—something mathematicians had failed to do.
The Benjamini-Hochberg procedure, a method cited over 130,000 times, was assumed to work reliably with correlated data, but the AI showed cases where it misses its target false discovery rate.
Although the practical gap is small, the result highlights AI's growing problem-solving capability in mathematics and raises questions about whether AI can reason to genuinely new knowledge.
What happened
A University of Pennsylvania statistics professor used OpenAI's GPT-5.6 Sol Pro to disprove a longstanding assumption about the Benjamini-Hochberg procedure, a widely-used statistical method for controlling false positives. The AI constructed a statistical model showing the method's false discovery rate can exceed its target level when data is correlated and normally distributed—something mathematicians had assumed but never proven in 30 years. GPT-5.6 Sol Pro completed the work in about 90 minutes; GPT-5.5 could not find a solution even after roughly 20 hours.
Why it matters
The Benjamini-Hochberg procedure, introduced in 1995, has been cited over 130,000 times and is used across modern statistics and scientific fields to filter out false alarms when testing thousands of hypotheses at once. While the gap the AI found is small in practical terms (0.104 versus the target 0.1), Berkeley statistician Will Fithian called the disproved conjecture 'the most interesting open problem in my area of statistics.' The result signals AI's advancing capability to solve problems that eluded human experts, marking what Fithian described as 'another marker of advancing AI capabilities whose consequences will reach far beyond math.'
What to watch
The AI's solution combined known statistical methods in an unusual way rather than inventing entirely new ones, leaving open the question of whether AI systems can generate genuinely novel knowledge or only recombine what they learned during training. Dobriban notes the result 'mainly matters for theory at this point' and that practical effects need further study. The full chat and code are publicly available.
Ask the AI about this article →
The Benjamini-Hochberg procedure has been a cornerstone of modern statistics since 1995, particularly in fields like genomics where researchers must filter false positives from thousands of simultaneous tests. For nearly three decades, the statistical community assumed the method would work reliably even when data points are correlated—a common real-world scenario, such as when genetic variants are inherited together. However, no one had formally proved this assumption. The University of Pennsylvania's Edgar Dobriban leveraged GPT-5.6 Sol Pro to fill that gap, constructing a counterexample showing that under certain conditions, the false discovery rate can exceed the target level. The speed of the solution is what stands out: 90 minutes versus roughly 20 hours for the previous generation model, and decades of unsuccessful human effort. Dobriban acknowledged that the AI's approach combined existing statistical methods in an unusual way rather than inventing new ones—a pattern seen in similar AI breakthroughs in mathematics. This raises a fundamental question about AI's reasoning: can these systems generate truly novel knowledge, or are they limited to sophisticated recombination of ideas learned during training? Even if recombination is the ceiling, Dobriban's work demonstrates practical value as a tool embedded in human workflows. The broader implications remain uncertain, though some researchers, including deep learning pioneer Richard Sutton, believe genuine self-improvement and generalization require capabilities beyond recombination.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
U.S. markets ended August higher, with the S&P 500 up 2.6% and the Nasdaq up 3.9%

Neurovia AI, an Abu Dhabi-based company, is pitching Saudi security agencies software that it says can compres…

AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…
