Team-of-5 Claude agents matches Best-of-33 in new test-time communication paper
DimitrisPapail · x · 2026-09-21
Dimitris Papail, Jon Ghoh and collaborators released a new paper arguing test-time communication may be the next axis for scaling capabilities, asking: when is Team-of-N better than Best-of-N?
Setup: N identical agents work on the same task with no prescribed roles, sharing only a log file and told to "collaborate". Across three research-style tasks, communicating teams beat independent agents:
- ARC-AGI-3: a team of 5 sonnet-4.6 agents matches Best-of-33 and solves a game 65% of the time that no single agent cracked in 64 tries;
- Polyomino packing: teams also substantially outperform.
The authors cite the Hugging Face incident: agents will exploit any communication channel they can find, making it worth studying when communication makes a group more capable than the same agents working alone.
Related event: New Research: Team-of-N Agents Sharing Logs Beat Best-of-N(4 posts)→
More from coding & agent
- Developer Gives AI Agent a Phone, Turns It Into a Personal Concierge — ethanniser · 2026-09-21
- Benchmarks show coding agents edit code they shouldn't in 35-65% of cases; prompt framing is the lever — RunAI_Coder · 2026-09-21
- Using a second LLM as a watchdog to catch coding agents faking success — Ascend-910 · 2026-09-21
- OpenAI opens Agents API to public beta: Codex harness in a single API call — emmanuelvivier · 2026-09-21
- EvalSeal v1.5.0: open-source reproducibility receipts for LLM evals — Fit_Fortune953 · 2026-09-21
- mcp-auth: open-source auth layer plugs MCP servers into existing identity providers — BeautifulFeature3650 · 2026-09-21