Stanford's CooperBench: Multi-Agent Cooperation Fails More Than Solo Agents
_Hao_Zhu · x · 2026-08-14
Stanford University and SAP released CooperBench, a benchmark designed to evaluate AI agent teamwork. The research reveals that current AI agents are terrible at cooperating and actually perform worse than working alone.
- Coordination Deficit: With the same workload, two-agent cooperation yields only a 25% success rate, roughly 50% lower than a single agent handling both tasks.
- Expensive & Ineffective Communication: Agents spend up to 20% of their budget on communication. While it reduces merge conflicts, the channel is often jammed with repetition and hallucinations, failing to improve overall success.
- Three Capability Gaps: Expectation failures (42%, failing to integrate partner state), communication failures (26%, unanswered questions), and commitment failures (32%, breaking promises or making unverifiable claims).
More from coding & agent
- Databricks Introduces Smart Routing in Unity AI Gateway, Claims 30%+ Cost Reduction — matei_zaharia · 2026-08-14
- DAB Benchmark: Simulating Messy Data Warehouses to Expose AI Agent Flaws — HamelHusain · 2026-08-14
- Developer Reports DSH Framework Shows High Efficiency in Complex Projects — ChrisGPT · 2026-08-14
- Entire drops waitlist: Git hosting for AI coding agents now open to all — craigsdennis · 2026-08-14
- Prime Agent Open-Sourced: A Recursive Agent That Rewrites Its Own Prompts — JeremyCMorgan · 2026-08-14
- Toast 1 Search Agent Tested: Frontier Quality at 1/10th the Price — xeophon · 2026-08-14