Berkeley paper: communicating agent teams match 4x more independent agents on ARC-AGI-3
DimitrisPapail · x · 2026-09-22
A UC Berkeley & Microsoft Research paper, Scaling Discovery through Test-Time Communication, tests whether agents sharing a log at inference time beat independent parallel attempts.
- N role-free agents worked on the same task via a shared text log; benchmark was ARC-AGI-3, which demands novel problem solving.
- team@k matches the success rate of 4k independent agents, and the gap grows with k — gains compound with scale.
- Not just efficiency: tasks no single agent can solve get solved reliably by the team.
- On research-style tasks: polyomino packing results beat best@k and the prior best-known score; on MNIST classifier compression, four agents produced a 1,957-byte classifier at 99.4% test accuracy, surpassing both the best human solution and best single-agent run.
- Caveat: with limited compute or no reliable progress verification, independent agents can still win.
Related event: Test-Time Communication: Agent Teams of 5 Match Best-of-33(20 posts)→
More from Research
- Frontend design framed as visual agent task with groupwise relative grading — stochasticchasm · 2026-09-22
- Team reportedly plans to open source 7,000 RL training environments — airesearch12 · 2026-09-22
- Why RL generalizes to reasoning but not literary writing, per AI researchers — phl43 · 2026-09-22
- Rethinking Policy Gradients: Score Centering Skips Importance Sampling Entirely — brandondamos · 2026-09-22
- Building one of the hardest on-policy lie datasets for Aletheia's Quest lie detection competition — hunarbatra · 2026-09-22
- Question's Gambit lifts deep research agents: GPT-5.5 hits 90.5% on BrowseComp-Plus — omarsar0 · 2026-09-22