Google’s 260-run study finds multi-agent systems can be 80.8% better — or 70% worse

alex_verem · x · 2026-07-25

Google Research, DeepMind, and MIT ran the largest controlled study on multi-agent systems, testing 260 configurations across OpenAI, Google, and Anthropic models on six benchmarks and five architectures with the same tools and compute budgets.

The key takeaway is task-dependent performance:

The paper's practical advice is simple: only go multi-agent when the work can truly be partitioned into independent subproblems.

Related event: Google's Massive Experiments Reveal Multi-Agent Systems' Double-Edged Sword(3 posts)→

Original post →

More from Research

Research channel →