AgentWorld Benchmark: Best Model Hits Only 52% on Multi-Agent Collaboration

新智元 · wechat · 2026-10-10

OpenAgents with Columbia, UPenn and others released AgentWorld, a benchmark where up to 10 LLM-driven agents with different roles, skills and resources collaborate in an MMORPG sandbox via in-game chat and actions. It includes 100 human-designed tasks plus 100 augmented variants, requiring 3-20 agents over dozens of rounds.

Three design choices expose collaboration failures: role asymmetry, black-box interaction (no access to teammates' internal reasoning), and multi-round execution. Under uniform settings, Gemini3 Flash scored 52% on main tasks, Claude Haiku 4.5 45%, GPT-5 Mini 36%, DeepSeek R1-70B 20%. Notably GPT-5 Mini sent 44 messages per task yet scored lowest, while Gemini sent just 11 and scored highest — message volume doesn't substitute for collaboration quality.

The paper also proposes CCE (Causal Collaboration Effectiveness), tracing which actions causally contributed to success. Failure analysis shows breakdowns in coordination details: repeated questions, wrong recipients, false item reports, and premature task-completion claims.

Original post →

More from coding & agent

coding & agent channel →