AIToday
Large Language ModelsOpen-Source AISiliconANGLE AIPublished: Oct 3, 2026, 04:00 JST

Ai2's Olmo-core 3 hits 52,000 tokens per second, topping Megatron-core

Ai2's Olmo-core 3 hits 52,000 tokens per second, topping Megatron-core

3 Key Points

  1. What happened

    The Allen Institute for AI announced Olmo-core 3 on Thursday, a training framework it says scaled mixture-of-experts models to over one trillion parameters while processing 52,000 tokens per second on Nvidia B3000 GPUs, about 2.7 times Nvidia's Megatron-core.

  2. Why it matters

    The framework is meant to lower the memory and networking costs that keep trillion-parameter models out of reach for researchers without state or enterprise infrastructure, which may widen who can train these systems.

  3. What to watch

    The throughput figure is a first-party benchmark, and the test is whether outside developers reproduce it on other hardware. The project and related systems are already on GitHub.

WHO IT HITSAI research teams without large-scale compute budgets are the clearest beneficiaries, since the framework is pitched at making trillion-parameter training feasible outside state and enterprise infrastructure. Cloud and GPU infrastructure teams may also weigh the throughput claim against their existing Megatron-core setups.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Mixture-of-experts models work differently from dense ones: they split computation across specialist portions of the model each time a token is generated, while dense models activate the whole model. That lets a MoE hold far more total parameters while computing with only a small number of them, cutting the compute per token. The catch, as Ai2 describes it, is that the full model still has to be stored across GPU memory, and coordinating networking between experts during training creates extra costs.

Olmo-core 3 is Ai2's answer to that catch. It uses expert parallelism to spread experts across multiple GPUs, so each card stores only part of the full expert pool. It also splits the model's layers across groups of GPUs and uses a distributed optimizer that spreads optimizer state across GPUs instead of keeping full copies on each one. Ai2 says this reduces memory overhead because the entire model and its training state need not be held in memory at once, and it adds support for MXFP8, a number format that represents some values with fewer bits to cut computation and data movement.

The stated ambition is about access: trillion-parameter models are often beyond the reach of those without state or enterprise infrastructure, and Ai2 says the framework would also let researchers adapt MoE training to different hardware and experiment with routing and parallelism. What the outcome hinges on is whether those gains hold outside Ai2's own benchmarks — especially given the comparison here is against Nvidia's Megatron-core on Nvidia's own GPUs. For researchers watching the space, the open GitHub release is the practical test.

FAQ
How much faster is Olmo-core 3 than Nvidia's Megatron-core?
Ai2 says Olmo-core 3 processed 52,000 tokens per second on Nvidia B3000 GPUs for a 47-billion parameter model, against roughly 19,400 tokens per second for Megatron-core — about 2.7 times the throughput.
Where can developers get Olmo-core 3?
The project and related systems are currently available for developers and the open-source community on GitHub.
How many experts does Olmo-core 3 let a model use?
Ai2 says the framework lets the expert pool grow from eight to 128 while still selecting only four experts per token.
SiliconANGLE AIRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI fires three safety researchers over leaks