Stanford paper: self-organizing agent teams hit 66.7% vs 48.8% for best single model

rohanpaul_ai · x · 2026-09-25

A Stanford + Together AI paper introduces Self-Organizing Agent Teams (SAT), rejecting the debate-and-vote multi-agent pattern. One model reviews past team chats and rewrites the team playbook (who checks whom, who plays devil's advocate), learning from just 15 math and 25 grad-level problems. Across five math/physics benchmarks, 3-model teams average 66.7% vs. 48.8% for the strongest member, 58.7% compute-matched, and 59.0% for a perfect router — beating even oracle routing on AIME 2026 by 13.4 points. Collaboration produced answers no member had.

Related event: Stanford's self-organizing agent teams beat best single models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →