AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Sep 11, 2026, 01:00 JST2 min read

Astra tops ErdosBench as OpenAI skips math on purpose

Astra tops ErdosBench as OpenAI skips math on purpose

3 Key Points

  1. What happened

    OpenAI's GPT-6 Astra took first place on ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open math problems (43 fully) and disproving 27. Chief scientist Jakub Pachocki says math optimization was deliberately deprioritized.

  2. Why it matters

    Benchmark developer Przemek Chojecki called it "a solid 5%-10% gain on various math-research skills tested." Astra's win came without targeted math optimization — evidence, the article argues, for a "spiky" AI trajectory where strength in one area means cutting back elsewhere.

  3. What to watch

    Since the essay was published, OpenAI has reportedly trained better internal math models, and RSI and AI safety topics have exploded — the tension between math progress and self-improvement priorities is unresolved. Watch whether Fable 5.1's 3.25 score holds, since Astra still solved the most problems overall.

WHO IT HITSAI research teams and mathematicians tracking model capabilities will need to read benchmark wins with more caution — a top score may signal a byproduct of other priorities rather than a targeted push. Math researchers face the practical question of who verifies proofs if models keep producing them faster than humans can check.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

OpenAI had put math wins front and center in its first announcement, which makes Pachocki's admission in "An Alien Mind" notable: the company believes it could make its models better at math research with additional focus, but chooses not to, citing urgency around recursive self-improvement and automated alignment research. That trade-off is the article's central evidence for what it calls the "spiky" thesis of AI development — extreme strength in select domains rather than broad, gradual improvement across all tasks.

Cambridge researcher Adam Hunt's visualization contrasts the "mainstream AGI" path of broad gradual improvement with an increasingly spiky trajectory, where coding and math surge while language quality, common sense, or social reasoning stagnate or regress. The article presents Pachocki's statement as evidence for the latter, noting that even the leading AI lab can't push maximum progress in every direction at once. Behind the RSI priority, however, is the hope that the model will eventually make those optimizations itself, scaling faster across the board — including in math.

Whether Astra is 5 or 50 percent better than its predecessor, the deeper question for mathematics remains: as Terence Tao raised at the 2026 International Congress of Mathematicians, if AI models keep producing proofs faster than humans can check them, the field risks shifting from proof scarcity to proof overload, and the critical task becomes deciding which results actually matter. The hardest problems remain unsolved for now, which may buy the discipline some time, but the direction still seems set — even if some mathematicians doubt language models can deliver real breakthroughs without human-like creativity.

FAQ
What is ErdosBench?
It is a math benchmark developed by Przemek Chojecki at ulam.ai, covering 226 open math problems inspired by the famous Erdős problems.
Why didn't OpenAI optimize Astra for math?
In his essay "An Alien Mind," chief scientist Jakub Pachocki writes that OpenAI could make models better at math research with additional focus but does not prioritize it because of urgency around recursive self-improvement and automated alignment research.
Has Astra been beaten on the benchmark?
Yes, Fable 5.1 has since knocked GPT-6 Astra off the top of the ErdosBench with a score of 3.25, edging out Astra's 3.23, though Astra still solved the most problems overall.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepSeek V4.1-Flash: 763B model beats V4 Pro on AA Index 40Latent Space · 38m ago
  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 6h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 6h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAnthropic's Jacob Coxon exits, warns of AI race