Self-organizing agent teams average 66.7% vs 48.8% for best single model, Stanford and Together AI paper finds
james_y_zou · x · 2026-09-24
- A paper from Stanford and Together AI shows letting agent teams learn their own collaboration strategy far outperforms expectations.
- A self-organizing team of three models (o3-mini, Claude Sonnet 4, DeepSeek-V3) averaged 66.7% across five math and physics benchmarks.
- Comparisons: the strongest member alone scored 48.8%, and even a perfect router over members' independent answers only hit 59.0%. On AIME 2026 the team reached 71.2%, 13.4 points above the router.
- Method: one member reviews the team's earlier exchanges and rewrites the teamwork strategy — roles, phase ordering, participation, and how partial answers are merged.
- The authors note research on effective agent collaboration/communication is lacking and self-organizing agents are more capable than previously thought.
More from coding & agent
- Claude Opus 5.5 Models, Textures and Rigs a Blender Character in Pure Code — Now a Gamedev Skill — TAbrodi · 2026-09-24
- A Ready-to-Use Prompt That Makes Your Agent Audit Its Own API Bills — gethackteam · 2026-09-24
- Claude Opus 5.5 One-Shots a 90s-Style Demoscene Demo in C/C++ and OpenGL — dreamwieber · 2026-09-24
- Months-long Claude-built Rummy 500 game opens free on web, rebuilt with Opus 5.5 — AIandDesign · 2026-09-24
- Garry Tan: agents are reshaping startup customer acquisition in both directions — garrytan · 2026-09-24
- Pattern MCP forces coding agents to reuse components and follow your design system — DonR954 · 2026-09-24