AIToday

LLMs tested on 2026 Math Olympiad problems; frontier models dominate

r/MachineLearning14h agoSend on LINE

Key takeaway

Researchers tested multiple large language models on unpublished 2026 International Mathematical Olympiad problems—a benchmark designed to measure general reasoning because the problems are new and mathematically complex. Frontier models (Claude Sonnet and Opus) scored near-perfectly, while open-weight and mid-tier models performed worse but improved significantly when paired with AutoFyn, a custom multi-agent workflow system the team built. The results show that frontier models have largely solved IMO-level reasoning, though specialized engineering can help smaller models narrow the gap.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Researchers benchmarked multiple large language models (AI systems that understand and generate text) against unpublished 2026 International Mathematical Olympiad problems. Frontier models (Claude Sonnet and Claude Opus, made by Anthropic) scored near-perfectly regardless of setup, while open-weight model GLM and mid-tier models performed worse but improved when paired with AutoFyn, a custom multi-agent orchestration system the team developed.

  • Why it matters

    IMO problems serve as a strong test of general reasoning ability because they are mathematically complex, multi-step tasks that have never appeared in any model's training data. The results show that current frontier models have largely solved this class of problem, while smaller or open-source models still lag—and that specialized harnesses (workflow tools) can meaningfully close the gap. For businesses evaluating which model to use for complex reasoning tasks, this suggests frontier models are necessary for top performance, though intermediate models can be improved with the right engineering.

  • What to watch

    The researchers used a different frontier model to grade all results. Numerical scores and full methodology are available in the attached paper.

In Depth

A group of researchers set out to evaluate how well different large language models perform on a meaningful test of mathematical reasoning: unpublished problems from the 2026 International Mathematical Olympiad (IMO). The motivation was straightforward: IMO problems are an exceptionally clean benchmark because they are new (so they cannot have appeared in training data), mathematically challenging, and representative of complex multi-step reasoning—exactly the kind of task where different models and different execution strategies might diverge.

The study compared several model categories. Frontier models—Anthropic's Claude Sonnet and Claude Opus—achieved near-perfect or perfect scores regardless of how they were set up or what supporting tools were used. Mid-tier and smaller models, by contrast, started much lower on the leaderboard. The open-weight model GLM (an openly available alternative to closed commercial models) performed roughly at the level of Sonnet without any special harness or orchestration, but still below the frontier tier.

Where the study becomes more instructive is in the intervention: researchers paired the weaker models with AutoFyn, a customizable multi-agent orchestration harness they had developed. This is not a language model itself, but rather a system that coordinates how a model tackles a problem—breaking it into steps, trying different approaches, and integrating intermediate results. The effect was consistent. Sonnet's performance improved further. GLM improved to levels comparable with Sonnet operating without harness. Even Opus, already at the frontier, showed gains when paired with AutoFyn.

Despite these improvements, no amount of harness engineering allowed non-frontier models to fully match frontier performance. The researchers noted that even with AutoFyn's support, they "were not able to match the performance of the frontier models." All results were graded by a different frontier model, ensuring consistent evaluation across all submissions. The full numerical scores and methodology are available in an attached paper, indicating this was a rigorous, reproducible benchmark exercise.

Context & Analysis

International Mathematical Olympiad problems have emerged as a meaningful test for large language models because they uniquely combine novelty (the 2026 problems had not been published when any of these models were trained), complexity (multi-step mathematical reasoning), and a clear right answer. This study exploited that by comparing how different model tiers—frontier systems versus mid-tier and open-weight alternatives—perform on unseen, genuinely hard problems.

The gap between frontier and non-frontier models is stark: Sonnet and Opus reached near-perfect or perfect scores regardless of how they were invoked or what supporting tools surrounded them. By contrast, smaller models like GLM and mid-tier systems started well below that threshold. However, the introduction of AutoFyn, a multi-agent orchestration harness, proved that engineering alone can partially bridge the gap. Sonnet and Opus improved further with AutoFyn, and GLM's performance with the harness rose to competitive levels with mid-tier models without harness support. This suggests that orchestration and workflow design matter, yet frontier model capability still remains a ceiling that smaller models cannot entirely overcome even with the right harness.

The grading itself was done by a frontier model, which is a methodological choice worth noting: the evaluation is consistent with how these systems are typically assessed, but introduces the question of whether a frontier model's scoring preference matches human or expert judgment on novel Olympiad problems.

FAQ

Which models were tested?
The study included frontier models (Sonnet and Opus), the open-weight model GLM, and other mid-tier systems. A different frontier model was used to grade all results.
What is AutoFyn?
AutoFyn is a customizable multi-agent harness (orchestration system) the researchers developed. It improved performance of non-frontier models when paired with them, though even with AutoFyn they could not match frontier model performance.
Why use IMO problems as a benchmark?
IMO problems are new and not in any model's training data, they are mathematically hard and proxy for general intelligence, and they are complex multi-step tasks that benefit from orchestration and harness engineering.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime