
What happened
Toby Ord's analysis of swarm scaling finds a 4-agent swarm needed about twice the total tokens for the same performance, but half the tokens per agent, so it could theoretically finish the task in half the time.
Why it matters
Swarms act as a new form of inference-scaling where parallel agents buy wall-clock speed, though returns diminish as agent counts grow — Ord compares this to economists' "stepping on toes" tax on coordinating large human groups.
What to watch
Ord says scaling agents 10x yields only 3x to 5x the performance of using 10x the tokens with one agent, and he notes he had hoped the value of λ for AI agents would be lower, which would make an intelligence explosion less likely.
WHO IT HITSAI lab researchers and product teams weighing whether to spend more tokens on parallel agents or on longer single-agent reasoning now have a concrete speed-versus-cost trade-off to use in planning.
Summaries like this, in your inbox every morning.
Ord frames AI swarms as a new form of inference-scaling, distinct from the two levers the field has mostly used so far: choosing the right mix of compute and data for a trained model, then spending inference budget on thinking through longer chains of thought and tool calls. Swarms add a third parameter — how many agents you run at once — and Ord's short analysis gives that parameter a shape. The speed benefit is real, but it comes with a token cost and a coordination penalty that he compares to the "stepping on toes" tax economists observe when large groups of people try to work together.
His most striking note is about what this implies for the odds of an intelligence explosion. Ord says he had hoped the value of λ for AI agents would be lower, because a lower λ would make an intelligence explosion less likely; the analysis instead suggests swarms could raise the chance of one. That reading matters for anyone tracking how fast AI capabilities could compound, though it rests on his modeling rather than on measured deployments.
A separate strand in the same newsletter points in a related direction. C5R Corp's SciUniverse benchmark tests how well AI systems operate a mostly automated lab, and Claude Fable 5.1 (xhigh) leads at a 45.3% pass rate with a cost-per-task of $40.61. DeepMind's paper on an Automated Scientific Economy argues the bottleneck for AI scientists will be physical resources and empirical validation rather than idea generation. Together with Ord's swarm analysis, these suggest the near-term story is less about raw model capability and more about how efficiently parallel agents, scarce lab equipment, and research budgets get allocated — a question whose answer will shape which labs and institutions can actually turn AI capability into results.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Ghost AI raised $11 million, led by Andreessen Horowitz with Abstract, Audacious Ventures, Nova and SV Angel…
OpenAI will roll out invisible text watermarks to all ChatGPT and Codex plan users in the EU within weeks, and…

Reflection announced Beam, a 501B-parameter open-weight model with 23B active parameters, claiming parity with…

A Stanford, Carnegie Mellon, UC Berkeley, and Microsoft Research team ran 6,800+ math, coding, and science tas…

Reflection AI launched Beam, a 501 billion-parameter open-source LLM
A step-by-step guide fine-tunes Muse Glimmer, Meta's 30B vision model, locally for equation-to-LaTeX conversio…
