
What happened
Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench, solo and as teams, at medium and maximum reasoning. Teams cost 1.8x to 5.1x more, and only one of four comparisons — GPT-6 Sol at medium reasoning, scoring 7.3 points higher — showed a statistically significant gain.
Why it matters
Since the teams consumed far more tokens for almost no measurable quality gain, the extra spend on multi-agent setups appears wasted when models already run at full compute.
WHO IT HITSTeams building AI coding assistants and other agentic workflows will want to check whether paying 1.8x to 5.1x more for agent teams actually lifts output quality, since Vals AI found it mostly does not.
Summaries like this, in your inbox every morning.
The finding lands against a backdrop of rising enthusiasm for multi-agent architectures, where teams of AI models are pitched as a way to tackle harder problems. Vals AI's benchmark, run on its Vibe Code Bench with GPT-6 Sol and Claude Opus 5.5, tested that premise directly by comparing each model solo and in teams at medium and maximum reasoning effort. The cost gap was wide — 1.8x to 5.1x — and the quality gap was mostly invisible, with only GPT-6 Sol at medium reasoning showing a statistically significant improvement of 7.3 points.
Anthropic's own data with Opus 5.5 points in the same direction. In its tests, larger teams reached a given performance level faster, but going from ten to 100 agents only nudged scores up slightly after 24 hours. On the knowledge base task, the score moved from 0.53 with one agent to 0.74 with 100 agents — a gain that came with far more tokens. Fable 5.1 showed stronger quality gains on the Lean theorem proving task above ten agents, yet still scored below Opus 5.5 across all tests, and its knowledge base score dipped slightly from 30 to 100 agents.
The practical caveat comes from OpenAI researcher Noam Brown, who told the Dwarkesh Podcast that the benefit depends heavily on the task: web research and math parallelize well, but writing a novel does not, and throwing 10,000 agents at a novel would be as pointless as throwing 10,000 people at it. OpenAI developer Eric Provencher has warned that agent swarms are most likely wasted money because coordination between agents breaks down — a cost he called the coordination tax.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
AMD closed at $608.10 on October 9, down 6.36% from its October 6 record close of $649.42, after a Financial T…

PolyAI CTO Shawn Wen told the HumanX conference that full-duplex models — which can speak while listening — no…

In a SiliconANGLE essay, Imagine Art's M
ルールベース株研 started 投資可否チェッカー on 2026年9月25日 and 急騰株チェッカー on 2026年10月8日, reaching over 300 commits and more than 8…

Tested on Windows 10 with Claude Code 2.1.296, a deny rule written as Edit(../shared/partner/**) did not stop…

A writer had Google Gemini read his health check results; Gemini warned of a future dialysis risk linked to hi…
