Stanford paper: self-organizing agent team hits 66.7%, beats best single model's 48.8%
dair_ai · x · 2026-09-23
A Stanford/Together AI paper shows why letting agent teams learn their own way of working together pays off:
- Three models (o3-mini, Claude Sonnet 4, DeepSeek-V3) as a self-organizing team averaged 66.7% across five math/physics benchmarks
- The strongest single member scored 48.8%; a perfect router over independent answers reached 59.0%
- On AIME 2026 the team hit 71.2%, 13.4 points above the router
Key mechanism: one member reviews past team exchanges and rewrites the teamwork strategy—roles, phase ordering, participation, answer merging. Strategies learned from just 15 AIME 2024 problems transferred unchanged to held-out problems and four new benchmarks.
Related event: Stanford Paper: Self-Organizing Agent Teams Learn to Reason Together(5 posts)→
More from coding & agent
- Everyone builds AI agents, almost nobody builds the harness around them — alex_verem · 2026-09-23
- Claude generates a stunning Three.js Japanese boat scene, and people can't believe it — nateliason · 2026-09-23
- The Agentic AI Storage Shock: enterprise agents turn storage into the next bottleneck — BenBajarin · 2026-09-23
- Claude Opus 5.5 impresses at building animated maps in hands-on test — DavidmComfort · 2026-09-23
- Self-host Firecrawl and give agents a Firecrawl skill for cleaner, cheaper web reading — gregmushen · 2026-09-23
- Daytona and LlamaIndex Host SF AI Builders Demo Night With 378 RSVPs — llama_index · 2026-09-23