
OpenAI's unreleased model Astra has solved ten previously open mathematics problems by consuming roughly $2,000 in computational tokens, with humans then formalizing the proofs in Lean.
This represents a substantial advancement in AI's capacity to tackle rigorous mathematics, a field where large language models historically lagged far behind human mathematicians.
What happened
OpenAI's internal version of Astra, its next major model, solved ten previously open mathematics problems. The model required roughly $2,000 worth of tokens (at Sol API rates) to find the solutions, which humans then prepared into manuscripts and formalized using Lean, a proof-verification system.
Why it matters
The result demonstrates a significant leap in AI's ability to handle rigorous mathematics—a domain where LLMs historically struggled. This capability may signal that frontier models are now approaching human-level reasoning on complex mathematical proofs, which has implications for research and technical problem-solving across industries.
What to watch
OpenAI is releasing narrations of the model's reasoning process alongside each solution, offering transparency into how Astra approaches mathematical problem-solving. The full extent of Astra's advantage over prior models like Fable and Sol in mathematics remains unclear.
OpenAI has disclosed that an internal version of Astra, described as its next major model, has solved ten previously open problems in mathematics. The model operates with remarkable efficiency: the total computational cost—measured in tokens consumed during inference—amounts to roughly $2,000 at current Sol API pricing.
The workflow for each solution follows a structured pipeline. First, Astra generates the mathematical argument. Second, human researchers prepare these raw outputs into formal manuscripts suitable for publication or presentation. Third, Astra formalizes each argument in Lean, a proof assistant that verifies correctness at the symbolic level. This formalization step is crucial: it ensures that the solutions are not merely plausible-sounding prose but mathematically airtight proofs. Lean certification is a gold standard in formal verification, so solutions that pass this step represent genuine mathematical advances, not speculative outputs.
OpenAI is also releasing narrations of the model's reasoning process for each of the ten solutions. This transparency move allows researchers and practitioners to inspect how Astra approaches mathematical problem-solving step by step.
The significance of this result lies in reversing a long-standing limitation. Large language models have historically performed poorly on mathematics, a gap that researchers and observers frequently cited as evidence of fundamental gaps in reasoning. Astra's ability to solve open problems suggests that capability gap is narrowing, at least at the frontier. However, the article does not specify whether Astra represents a dramatic leap beyond prior OpenAI models like Fable and Sol, only that "we do know that Astra can do math. As in real math."
For years, mathematics has been a notable weak point for large language models. The ability to reliably solve complex proofs and discover new mathematical results has eluded even the most capable models. The announcement that Astra—OpenAI's next major model—has solved ten open mathematics problems marks a turning point in that trajectory.
The approach itself is instructive. Rather than simply outputting an answer, Astra's solutions were verified through formalization in Lean, a proof assistant that checks mathematical correctness mechanistically. This suggests the model is not merely pattern-matching or hallucinating plausible-sounding statements, but producing arguments rigorous enough to survive formal verification. The fact that humans needed to intervene to prepare manuscripts and formalize proofs indicates the workflow still requires human expertise, but the core mathematical reasoning—the part that historically stumbled—now appears to be within reach for the model.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Supermicro reported a sharp recovery in gross margins and issued fiscal 2027 revenue guidance that exceeded Wa…

IBM and Together AI announced a collaboration to deploy a dedicated Nvidia-powered inference cluster on IBM Cl…

Siliconware Precision Industries (SPIL), a subsidiary of ASE Technology Holding, held a groundbreaking ceremon…

CoreWeave reported second-quarter fiscal 2026 revenue of US$2.6 billion, up 112% year-over-year and 24% sequen…

CoreWeave, an AI cloud provider, more than doubled its second-quarter revenue while managing a $104 billion ba…

Palantir Technologies reported second-quarter revenue up 93% year over year to $1.94 billion on Aug

The AI news that matters, in one minute each morning.
Sign up free