
Leading mathematicians including Timothy Gowers and Peter Sarnak have concluded that large language models possess strong computational and combinatorial skills but cannot generate genuinely novel mathematical ideas.
The researchers attribute this limitation to LLMs' inability to develop new foundational abstractions from first principles, suggesting their recent performance gains reflect refinement in familiar domains rather than true creative reasoning.
This finding raises questions about whether LLMs are becoming genuinely more versatile or simply excelling within established problem spaces.
What happened
Mathematicians Timothy Gowers and Peter Sarnak, alongside DeepMind researcher Tom Zahavy, have concluded that large language models are skilled at combining existing mathematical methods but cannot generate genuinely novel ideas. Gowers notes that LLMs are good at trying many search paths but lack the intuition to identify productive routes in a vast search space. Sarnak emphasizes that AI can derive results from existing theory but fails to develop the foundational abstractions that underpin major proofs when starting from an elementary question.
Why it matters
The findings highlight a fundamental boundary in what current LLMs can and cannot do. Rather than representing a versatility breakthrough, their improving performance may reflect refinement in familiar problem spaces and benchmarks rather than true creative thinking. For mathematicians and researchers relying on AI for discovery, this suggests LLMs remain tools for execution rather than innovation.
What to watch
Zahavy's research identifies "manipulative abduction"—the ability to invent new foundational assumptions with no linguistic precedent—as the critical bottleneck. The assessment suggests that world models could offer a path forward, making this an emerging area of focus for advancing AI reasoning beyond its current limits.
Ask the AI about this article →
The critique from Gowers, Sarnak, and Zahavy points to a structural limitation in how current LLMs approach problem-solving. While these models excel at combining known methods and exploring multiple solution paths, they operate within an existing conceptual framework—they can execute but not reimagine the foundations upon which mathematics rests. The bottleneck identified by Zahavy, the inability to perform "manipulative abduction," reveals that LLMs lack a faculty that is central to mathematical creativity: the capacity to posit entirely new assumptions or abstractions when existing ones prove insufficient. This distinction matters because it clarifies what "improvement" in LLM performance actually means. Higher benchmark scores do not necessarily indicate that models have crossed into creative reasoning; they may simply show that models have become more fluent within the domains on which they were trained and tested. The suggestion that world models could offer a solution points to a direction for future research—one that moves beyond the current token-based architecture toward something closer to explicit causal or structural reasoning.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic plans to "match or beat" the size of SpaceX's $75 billion IPO (or $86.2 billion including the over-a…

Pew Research released a study on Thursday finding that over one-third (35%) of English-language web pages publ…

The article argues that non-expert managers and consultants—people whose only exposure to AI comes from ChatGP…

OpenAI is previewing transparent background support for GPT-Image-2 through its API, allowing users to generat…

OpenAI's GPT-5.6 Sol, launched July 9, drove a 35 percent revenue increase this quarter, with enterprise reven…

At the "AI on Chips: Semiconductor Industry Trends Forum" hosted by DIGITIMES, industry experts highlighted th…
