AgentWorld Benchmark: Best LLM Only Hits 52% on Long-Horizon Multi-Agent Collaboration
Raphael Shu · hf · 2026-09-28
AgentWorld is an open-source benchmark for evaluating long-horizon multi-agent collaboration among LLMs, and current models fare poorly.
- Design: 100 human-annotated tasks (plus 100 augmented variants) set in an MMORPG sandbox, requiring 3-20 agents with asymmetric roles to coordinate over 50+ interaction rounds via communication, joint planning, and resource sharing—under a blackbox setup where agents can't see each other's internal states.
- New metric: Causal Collaboration Effectiveness (CCE), a graph-based metric tracing causal dependencies between agent actions to measure what fraction of team effort actually contributed to outcomes, beyond binary task success.
- Results: Testing Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B, the best model reached only 52.0% task success, with systematic failures in communication breakdowns, role confusion, and inability to maintain shared plans.
More from coding & agent
- Study of 2,170 GitHub projects maps how the fast, low-cost Jev decision model is used in the wild — CUHK-CSE · 2026-09-28
- Anthropic releases free 37-minute guide on building AI agents that automate entire businesses — Aiden_Tech_Ai · 2026-09-28
- Spotify paper: synthetic multi-turn dialogs and self-improvement loops boost conversational recsys quality by 8% — _reachsumit · 2026-09-28
- RecToolBench: 1,200+ Task Benchmark Tests Recommender Agents on MCP Tool Orchestration — _reachsumit · 2026-09-28
- Grok Build v1.0.43 fixes MCP prompts going missing and hanging silently in minimal mode — XFreeze · 2026-09-28
- Dev Open-Sources Claude Code Skill That Makes Full AI Music Videos End-to-End — TimothyDuignan · 2026-09-28