Self-organizing agent teams hit 66.7% on math benchmarks vs 48.8% for their strongest member
zainhas · x · 2026-10-12
A Stanford-affiliated paper (arXiv:2609.22682) introduces Self-Organizing Agent Teams (SAT): fixed agent teams that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow — no preset task decomposition or routing.
Key findings
- Introduces "collaborative computation": agents exchange, challenge, repair, and synthesize partial reasoning into solutions no single member could produce alone.
- Teamwork strategies learned from just 15 math + 25 graduate-level knowledge problems transfer unchanged to unseen benchmarks.
- Average accuracy of 66.7% across five math/physics benchmarks vs 48.8% for the strongest member, 58.7% for compute-matched inference, and 59.0% for a perfect router over independent answers.
- On AIME 2026, SAT beats the perfect router by 13.4 points.
- The paper analyzes when self-organizing collaboration helps, pointing to demonstrability as a factor.
Takeaway: how a team organizes itself is itself a learnable capability, beyond fixed protocols and routing.
More from coding & agent
- PhD student uses nightly Grok agent to auto-fill Zotero with papers and open-access PDFs — Scobleizer · 2026-10-12
- Vibecoded Pokemon-NFL simulator games go viral, one-shot by NFL Twitter — nrehiew_ · 2026-10-12
- rea: agent-powered reverse engineering tool hits 99.2k GitHub stars — adam_dorr · 2026-10-12
- Santa Fe's David Krakauer framing: vibe coders are capable, not competent — so are LLMs — MarcJSchmidt · 2026-10-12
- 'Clean code is dead': devs rethink maintainability as agents write for machines, not humans — bendee983 · 2026-10-12
- Zed CEO: Agents Will Use Existing Tools, Not Code Desktop Apps From Scratch — zeeg · 2026-10-12