Toby Ord quantifies swarm scaling: 16x parallel sampling equals just 4.9x longer chain-of-thought

tobyordoxford · x · 2026-09-22

Philosopher Toby Ord used benchmark data published on OpenAI's GPT-5.6 launch page to quantitatively compare two scaling paths: widening parallel sampling (swarm scaling) versus lengthening the chain of thought.

The takeaway: per unit of compute, longer reasoning chains currently buy more performance than fanning out parallel samples.

Related event: Toby Ord quantifies swarm scaling: far less efficient than longer chain-of-thought(17 posts)→

Original post →

More from Research

Research channel →