Parallel beats serial at equal budget: best-of-N@T far outperforms best-of-1@N*T
ShangyinT · x · 2026-09-22
In a debate on multi-agent baselines, generatorman argues that the fair baseline for team-of-N with total token budget T is not best-of-N@T but best-of-1@NT — the same budget spent serially on one agent.
ShangyinT reports that for reasonably large NT, best-of-1@NT is much worse than best-of-N@T, i.e. parallel sampling/agents beat stacking the budget serially. DimitrisPapail confirms this is measurably true across all his non-science projects.
More from coding & agent
- 53 harness bugs: an engineering team's field notes on building eval harnesses — DrDatta_AIIMS · 2026-09-22
- atuin 18.23 ships infinite searchable scrollback — shell history now keeps command output — aronchick · 2026-09-22
- Open-source coding-agent skills for Google ADK on Google Cloud, distilled from real engineering lessons — Difficult_Design6676 · 2026-09-22
- After trimming his Claude Code harness, a Max 20x user now has 9 spare hours of quota a week — carlito_17 · 2026-09-22
- OpenAI shows how a surgeon uses Codex to build tools and review PubMed literature — OpenAIDevs · 2026-09-22
- Similarity Isn't Relevance: Four Layers Every Personal Agent Memory System Must Separate — sujingshen · 2026-09-22