New LLM Multi-Agent Coordination Benchmark: Most Models Struggle
ktessera · reddit · 2026-07-14
Researchers introduced a novel benchmark to evaluate LLMs' multi-agent coordination abilities in long-horizon, open-ended environments. In this setting, agents must collaborate on complex tasks like exploration, communication, resource trading, tool crafting, construction, and combat.
Key Findings:
- Overall Poor Performance: The 13 tested mainstream LLMs achieved an average normalized return of only about 6%.
- Gemini Stands Out: Under the most difficult settings, zero-shot Gemini 3.1 Pro performed exceptionally, rivaling top MARL (Multi-Agent Reinforcement Learning) agents trained for 1 billion steps.
- Coordination is the Bottleneck: Beyond long-horizon execution, inter-agent communication is a critical bottleneck. Ablation studies show that communication mechanisms impact coordination the most.
The project is open-source with interactive trajectory demos available.
Related event: Studies Highlight Deficiencies in LLM Multi-Agent Collaboration(6 posts)→
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22