Swarm scaling needs squared inference to match chain-of-thought gains, analysis of OpenAI curves finds

tobyordoxford · x · 2026-09-22

tobyord Oxford analyzes OpenAI's multi-agent swarm scaling curves and finds that spending extra inference on swarm size yields only about half the score increase of longer chain-of-thought — you need to square the inference scale-up of swarm size to match a scale-up of CoT. Since OpenAI ran no clean ablation holding CoT fixed while scaling agent count, he infers it by joining points with equal tokens-per-agent across the curves. A useful empirical signal for whether inference compute should go to CoT length or more agents.

Related event: Oxford's Toby Ord Quantifies Swarm Scaling: Agent Clusters Far Less Efficient Than Longer Chains of Thought(10 posts)→

Original post →

More from Models

Models channel →