New Benchmark for LLM Multi-Agent Collaboration
j_foerst · x · 2026-07-14
This post introduces a new multi-agent benchmark evaluating how 13 modern LLMs collaborate in long-horizon, open-ended worlds to explore, communicate, trade resources, craft tools, build structures, and combat monsters.
Key findings: Most agents performed poorly, averaging only about 6% normalized return. However, under the most difficult settings, zero-shot Gemini 3.1 Pro matched the performance of an optimal MARL agent trained for 1 billion environment steps.
The authors conclude that coordination capability is an independent bottleneck, beyond just "the ability to complete long tasks." Ablation studies showed that communication had the greatest impact on results.
Related event: Studies Highlight Deficiencies in LLM Multi-Agent Collaboration(6 posts)→
More from Models
- Daily AI brief: GPT-Live-1 in API, OpenAI pauses $200 Pro signups amid Astra demand — koltregaskes · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11