Researchers have published a new benchmark that evaluates how well 13 modern large language models can coordinate as agents in open-ended, long-horizon worlds.
Most agents performed poorly, averaging only ~6% normalised return, revealing that coordination is a distinct challenge beyond individual task competence.
Notably, zero-shot Gemini 3.1 Pro performed as well as the best multi-agent reinforcement learning agent trained for 1 billion environment steps on the hardest setting, and ablation tests showed that communication has the largest effect on coordination success.
What happened
Researchers evaluated 13 modern LLMs in a new benchmark requiring agents to coordinate in open-ended worlds—exploring, communicating, trading resources, crafting tools, building structures, and fighting mobs. Most agents averaged only ~6% normalised return, though zero-shot Gemini 3.1 Pro matched the best multi-agent reinforcement learning (MARL) agent trained for 1 billion environment steps on the hardest setting.
Why it matters
The benchmark reveals that coordination is a distinct bottleneck separate from long-horizon task competence. Communication emerged as the single largest factor in ablation tests, suggesting that how well agents exchange information—not just individual task ability—determines team success. This finding matters for understanding what language models still need to master before reliably working together in complex environments.
What to watch
The full paper, project page with leaderboard, code, and interactive traces are publicly available, allowing the research community to test their own agents and track progress on this coordination challenge.
Ask the AI about this article →
The benchmark addresses a practical gap in AI evaluation: while individual LLMs have demonstrated impressive capabilities on long-horizon reasoning tasks, far less is known about their ability to work together effectively in shared environments. By designing scenarios that require exploration, resource trading, tool crafting, and collective problem-solving, the researchers created a setting where agent ability must be paired with team coordination.
The striking performance gap—most agents at ~6% normalised return—suggests that modern LLMs out of the box lack robust multi-agent coordination strategies. The fact that zero-shot Gemini 3.1 Pro matched MARL agents trained for 1 billion environment steps on the hardest setting indicates that some foundation models have learned coordination behaviors, but the majority have not. The ablation finding that communication is the largest lever is concrete and actionable: it points toward communication protocols and information exchange as the most fruitful area for improving team performance, rather than raw individual capability.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Sonos announced a new app update with generative AI features, a new soundbar called the Beam Ultra, and its se…
