100 Claude agents no better than 10: Ramez Naam shows agent scaling tops out fast
zsakib_ · x · 2026-10-11
Planetary VC founder Ramez Naam shared hard numbers on multi-agent scaling, calling it "the worst scaling we've ever seen of anything":
- A single Claude Opus agent succeeds 70% of the time; going to 4 agents adds 3 points, 16 agents adds 2 more, and beyond 10 agents performance flatlines entirely — 100 agents gain nothing.
- He ranks scaling efficiency: pre-training > test-time compute > agent swarms. Peak single-entity intelligence matters far more than many lower-level collaborators, analogized to Columbus, Ohio's population failing to out-Einstein Einstein.
- The same interview argues the intelligence explosion isn't close: the AI self-improvement loop is 5-10x too weak. OpenAI researchers used 124x more tokens per person but ran only 1.6x more experiments — labs' public claims don't match their own data.
Related event: VC Founder's Tests Show Multi-Agent Scaling Far Below Expectations(2 posts)→
More from coding & agent
- Evals are replays: store tasks in R2, sandbox in Docker, and 90% of the work is measuring — danshipper · 2026-10-11
- Solo dev dilemma: self-hosted n8n or FastAPI on Cloud Run for LLM automations? — Mysterious_Profit696 · 2026-10-11
- Autoresearch agents get stuck in 'idea basins' — fork-and-flush offers a fix — menhguin · 2026-10-11
- JetBrains Air hands-on: running two parallel coding agents inside WebStorm — mhdfaran · 2026-10-11
- Vibe coding an Ori-style platformer with a giant boss chase using Opus 5.5 — chongdashu · 2026-10-11
- Dev builds tiny logger to find what eats OpenAI budget — token counts pointed to the wrong feature — Atm1n9 · 2026-10-11