AgentWorld Benchmark: Best Model Hits Only 52% on Multi-Agent Collaboration
新智元 · wechat · 2026-10-10
OpenAgents with Columbia, UPenn and others released AgentWorld, a benchmark where up to 10 LLM-driven agents with different roles, skills and resources collaborate in an MMORPG sandbox via in-game chat and actions. It includes 100 human-designed tasks plus 100 augmented variants, requiring 3-20 agents over dozens of rounds.
Three design choices expose collaboration failures: role asymmetry, black-box interaction (no access to teammates' internal reasoning), and multi-round execution. Under uniform settings, Gemini3 Flash scored 52% on main tasks, Claude Haiku 4.5 45%, GPT-5 Mini 36%, DeepSeek R1-70B 20%. Notably GPT-5 Mini sent 44 messages per task yet scored lowest, while Gemini sent just 11 and scored highest — message volume doesn't substitute for collaboration quality.
The paper also proposes CCE (Causal Collaboration Effectiveness), tracing which actions causally contributed to success. Failure analysis shows breakdowns in coordination details: repeated questions, wrong recipients, false item reports, and premature task-completion claims.
More from coding & agent
- Debugging 9K lines of 100% AI-generated code: standard tactics failed to find root cause — blaizedsouza · 2026-10-11
- ManimGX open-sources a Rust+wgpu 3D animation engine for agents, 90× faster than ManimCE — Scobleizer · 2026-10-11
- Redot Engine welcomes AI-generated PRs — but only if contributors understand their own code — esrtweet · 2026-10-11
- $500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context — Ok-Breadfruit-3523 · 2026-10-11
- Jev founder's harness guide: coding agents up to 200x faster, 400x cheaper — blaizedsouza · 2026-10-11
- First 24 hours with Hark agent: auto-connects accounts, fixes its own Notion errors, 2FA remains a hurdle — Scobleizer · 2026-10-11