Stanford/Together AI paper: agent teams hit 66.7% vs 48.8% for single agents
mark_k · x · 2026-09-27
A new paper from Stanford and Together AI shows AI agent teams learning to collaborate rather than follow fixed scripts.
- Agent teams learn when to divide roles, challenge each other, and merge partial solutions
- Across five math and physics benchmarks, teams averaged 66.7% accuracy vs 48.8% for their strongest member alone, and 58.7% when that member got the same compute budget
- Notably, teams even beat an oracle-like system that could pick the correct answer whenever any member found it independently — suggesting discussions produced answers no member reached alone
The authors note results still await independent replication.
More from coding & agent
- IBM open-sources Docling, a free Python library that converts any document to data — mdancho84 · 2026-09-27
- Dev reflects: coding now feels like a waste of time when LLMs solve 99% of problems — justalexoki · 2026-09-27
- Reward hacking bugs revealed: unpruned git history let models peek at patches, edit tests in shared sandbox — willcb · 2026-09-27
- A Content Pinball Machine built entirely with Claude Opus 5.5 satirizes viral randomness — CurieuxExplorer · 2026-09-27
- Dev caches hide 264 CVE-laden packages and 71.7 GiB no one audits, warns Cache Commander dev — julsimon · 2026-09-27
- WebMCP could become the HTML/API layer of the agentic web — Thionne_WTZ · 2026-09-27