Team-of-5 Claude agents matches Best-of-33 on ARC-AGI-3, solving one game 65% of the time
DimitrisPapail · x · 2026-09-21
Greg Kamradt (ARC Prize) highlighted a new paper by Dimitris Papail et al.: on ARC-AGI-3, a team of 5 sonnet-4.6 agents collaborating matches Best-of-33 independent agents, solving one game 65% of the time that no single agent cracked in 64 tries.
The method: N identical agents with no prescribed roles collaborate only through a shared text log. Communicating teams substantially beat independent agents across research-style tasks, supporting the authors' claim that test-time communication is a next axis for scaling capabilities.
More from coding & agent
- Exa MCP hits 5,000 GitHub stars as AI agents flock to its search integration — TheIshanGoswami · 2026-09-22
- 670,000 agent skills, no trust layer: bot scan finds 69% never reliably fire — markjeffrey · 2026-09-22
- Training on production traces: single-trajectory RL may unlock continual learning — rhythmrg · 2026-09-22
- Anthropic's Swiss cheese model explains why passing evals isn't enough for agents — hugobowne · 2026-09-22
- OpenAI's artists are now all using Codex in their workflow — andrew_n_carr · 2026-09-22
- TinyTorch: PyTorch's free curriculum to build an ML framework from scratch in 20 modules — PyTorch · 2026-09-22