
What happened
OpenAI's GPT-6 Astra took first place on ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open math problems (43 fully) and disproving 27. Chief scientist Jakub Pachocki says math optimization was deliberately deprioritized.
Why it matters
Benchmark developer Przemek Chojecki called it "a solid 5%-10% gain on various math-research skills tested." Astra's win came without targeted math optimization — evidence, the article argues, for a "spiky" AI trajectory where strength in one area means cutting back elsewhere.
What to watch
Since the essay was published, OpenAI has reportedly trained better internal math models, and RSI and AI safety topics have exploded — the tension between math progress and self-improvement priorities is unresolved. Watch whether Fable 5.1's 3.25 score holds, since Astra still solved the most problems overall.
WHO IT HITSAI research teams and mathematicians tracking model capabilities will need to read benchmark wins with more caution — a top score may signal a byproduct of other priorities rather than a targeted push. Math researchers face the practical question of who verifies proofs if models keep producing them faster than humans can check.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
OpenAI had put math wins front and center in its first announcement, which makes Pachocki's admission in "An Alien Mind" notable: the company believes it could make its models better at math research with additional focus, but chooses not to, citing urgency around recursive self-improvement and automated alignment research. That trade-off is the article's central evidence for what it calls the "spiky" thesis of AI development — extreme strength in select domains rather than broad, gradual improvement across all tasks.
Cambridge researcher Adam Hunt's visualization contrasts the "mainstream AGI" path of broad gradual improvement with an increasingly spiky trajectory, where coding and math surge while language quality, common sense, or social reasoning stagnate or regress. The article presents Pachocki's statement as evidence for the latter, noting that even the leading AI lab can't push maximum progress in every direction at once. Behind the RSI priority, however, is the hope that the model will eventually make those optimizations itself, scaling faster across the board — including in math.
Whether Astra is 5 or 50 percent better than its predecessor, the deeper question for mathematics remains: as Terence Tao raised at the 2026 International Congress of Mathematicians, if AI models keep producing proofs faster than humans can check them, the field risks shifting from proof scarcity to proof overload, and the critical task becomes deciding which results actually matter. The hardest problems remain unsolved for now, which may buy the discipline some time, but the direction still seems set — even if some mathematicians doubt language models can deliver real breakthroughs without human-like creativity.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
DeepSeek launched V4.1-Flash, a 763B-parameter open-weight model with a causal encoder-decoder architecture

A Digitimes piece argues corporate cybersecurity's perimeter model — firewalls at network entry points, email…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…

Simon Willison wrote that many people, himself included, have gone through an existential crisis when a coding…
