AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Aug 17, 2026, 01:00 JST2 min read

Top mathematicians: LLMs excel at calculation, fail at creative proof

Top mathematicians: LLMs excel at calculation, fail at creative proof

Key takeaway

  • Leading mathematicians including Timothy Gowers and Peter Sarnak have concluded that large language models possess strong computational and combinatorial skills but cannot generate genuinely novel mathematical ideas.

  • The researchers attribute this limitation to LLMs' inability to develop new foundational abstractions from first principles, suggesting their recent performance gains reflect refinement in familiar domains rather than true creative reasoning.

  • This finding raises questions about whether LLMs are becoming genuinely more versatile or simply excelling within established problem spaces.

3 Key Points

  1. What happened

    Mathematicians Timothy Gowers and Peter Sarnak, alongside DeepMind researcher Tom Zahavy, have concluded that large language models are skilled at combining existing mathematical methods but cannot generate genuinely novel ideas. Gowers notes that LLMs are good at trying many search paths but lack the intuition to identify productive routes in a vast search space. Sarnak emphasizes that AI can derive results from existing theory but fails to develop the foundational abstractions that underpin major proofs when starting from an elementary question.

  2. Why it matters

    The findings highlight a fundamental boundary in what current LLMs can and cannot do. Rather than representing a versatility breakthrough, their improving performance may reflect refinement in familiar problem spaces and benchmarks rather than true creative thinking. For mathematicians and researchers relying on AI for discovery, this suggests LLMs remain tools for execution rather than innovation.

  3. What to watch

    Zahavy's research identifies "manipulative abduction"—the ability to invent new foundational assumptions with no linguistic precedent—as the critical bottleneck. The assessment suggests that world models could offer a path forward, making this an emerging area of focus for advancing AI reasoning beyond its current limits.

Ask the AI about this article →

Context & Analysis

The critique from Gowers, Sarnak, and Zahavy points to a structural limitation in how current LLMs approach problem-solving. While these models excel at combining known methods and exploring multiple solution paths, they operate within an existing conceptual framework—they can execute but not reimagine the foundations upon which mathematics rests. The bottleneck identified by Zahavy, the inability to perform "manipulative abduction," reveals that LLMs lack a faculty that is central to mathematical creativity: the capacity to posit entirely new assumptions or abstractions when existing ones prove insufficient. This distinction matters because it clarifies what "improvement" in LLM performance actually means. Higher benchmark scores do not necessarily indicate that models have crossed into creative reasoning; they may simply show that models have become more fluent within the domains on which they were trained and tested. The suggestion that world models could offer a solution points to a direction for future research—one that moves beyond the current token-based architecture toward something closer to explicit causal or structural reasoning.

FAQ

What specific weakness do mathematicians identify in LLMs?
Gowers and Sarnak argue that LLMs lack intuition for selecting productive search paths from a vast space and cannot develop the foundational abstractions needed for major proofs when starting from elementary questions. DeepMind's Tom Zahavy pinpoints "manipulative abduction"—the ability to invent new foundational assumptions with no linguistic precedent—as the key missing capability.
What do these findings suggest about LLM improvement?
The assessments feed into a broader debate about whether LLMs are actually becoming more versatile or simply getting better at benchmarks and familiar problem spaces. The mathematicians' conclusions suggest the latter interpretation may be more accurate.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCloudflare announces Kitesurf, a lightweight headless browser for AI agents

The AI news that matters, in one minute each morning.

Sign up free