Rule of thumb for sub-agents: ~50% latency gain per agent at 100% extra token cost

pvncher · x · 2026-10-09

A practical take on multi-agent architectures: each extra sub-agent buys roughly 50% latency improvement at 100% more token cost, so blasting sub-agents on every task just burns usage limits. Sub-agents make sense for fanning out context gathering across channels (Notion, Slack) to smaller models, or orchestrating discrete, independent, well-planned tasks. Small models offset cost but think and work slower—sometimes fast mode is simply cheaper.

Related event: Anthropic Engineer's Rule of Thumb: Each Sub-Agent Cuts Latency 50%, Doubles Tokens(3 posts)→

Original post →

More from coding & agent

coding & agent channel →