AIToday
Large Language ModelsAI Coding AssistantsTHE DECODERPublished: Oct 12, 2026, 01:00 JST

Vals AI: agent teams cost 5.1x more, barely beat solo agents

Vals AI: agent teams cost 5.1x more, barely beat solo agents

3 Key Points

  1. What happened

    Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench, solo and as teams, at medium and maximum reasoning. Teams cost 1.8x to 5.1x more, and only one of four comparisons — GPT-6 Sol at medium reasoning, scoring 7.3 points higher — showed a statistically significant gain.

  2. Why it matters

    Since the teams consumed far more tokens for almost no measurable quality gain, the extra spend on multi-agent setups appears wasted when models already run at full compute.

WHO IT HITSTeams building AI coding assistants and other agentic workflows will want to check whether paying 1.8x to 5.1x more for agent teams actually lifts output quality, since Vals AI found it mostly does not.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The finding lands against a backdrop of rising enthusiasm for multi-agent architectures, where teams of AI models are pitched as a way to tackle harder problems. Vals AI's benchmark, run on its Vibe Code Bench with GPT-6 Sol and Claude Opus 5.5, tested that premise directly by comparing each model solo and in teams at medium and maximum reasoning effort. The cost gap was wide — 1.8x to 5.1x — and the quality gap was mostly invisible, with only GPT-6 Sol at medium reasoning showing a statistically significant improvement of 7.3 points.

Anthropic's own data with Opus 5.5 points in the same direction. In its tests, larger teams reached a given performance level faster, but going from ten to 100 agents only nudged scores up slightly after 24 hours. On the knowledge base task, the score moved from 0.53 with one agent to 0.74 with 100 agents — a gain that came with far more tokens. Fable 5.1 showed stronger quality gains on the Lean theorem proving task above ten agents, yet still scored below Opus 5.5 across all tests, and its knowledge base score dipped slightly from 30 to 100 agents.

The practical caveat comes from OpenAI researcher Noam Brown, who told the Dwarkesh Podcast that the benefit depends heavily on the task: web research and math parallelize well, but writing a novel does not, and throwing 10,000 agents at a novel would be as pointless as throwing 10,000 people at it. OpenAI developer Eric Provencher has warned that agent swarms are most likely wasted money because coordination between agents breaks down — a cost he called the coordination tax.

FAQ
How much more do AI agent teams cost than single agents?
According to Vals AI's tests, agent teams cost between 1.8x and 5.1x more than single agents.
When did agent teams actually improve results in the tests?
Only one of four comparisons showed a statistically significant gain: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher.
What did OpenAI's Noam Brown say about multi-agent systems?
He said on the Dwarkesh Podcast that multi-agent systems mainly buy speed, not better quality — four agents solved tasks twice as fast but cost twice as much.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articlePolyAI CTO Shawn Wen: voice AI lacks its "ChatGPT moment"