Million-dollar AI swarms: progress will come from building 'unit tests' outside the models
mentalgeorge · x · 2026-09-17
Six thoughts on million-dollar AI 'swarms' as a distinct workload type:
- Test-time scaling is jagged: marginal-token dose-response varies hugely across distributions
- Time-horizons rise but diverge: HuggingFace hack or Navier-Stokes far exceeds booking flights
- Not all tasks suit swarming — ones with built-in verification (math, cyber) do best
- Near-term swarm significance depends on how many tasks have built-in verification
- Much AI progress will happen outside models: building 'unit test' equivalents for non-coding problems
- Picking problems to point swarms at is itself hard — OpenAI only tried Navier-Stokes due to Twitter rumors
Reply from 1a3orn: swarms may be the future as the only way to get 10,000x more compute at use-time than train-time.
Related event: Researchers Debate AI Swarms as Path to 10,000x Inference Compute(4 posts)→
More from AGI Musings
- Alignment researcher rebuts David Sacks: tail risks during RSI aren't an engineering problem — dhadfieldmenell · 2026-09-17
- Aza Raskin on the Utopias Podcast: can we still make humane technology? — aza · 2026-09-17
- Noah Giansiracusa guest posts on Terry Tao's blog: let the diners into the kitchen — AlexKontorovich · 2026-09-17
- In Dec 2024, AI Researchers Didn't Expect a Millennium Problem Solved Until 2054 — Tolopono · 2026-09-17
- Researchers debate AI risk framing: instant-apocalypse narratives vs. boiling-frog gradual harm — eigenhector · 2026-09-17
- Is domain-specific training a constant-factor gain or a scaling exponent change? — eigenron · 2026-09-17