Small model swarms are a mirage: 20x cheaper models burn 40-50x more tokens
pvncher · x · 2026-10-08
pvncher argues the appeal of swarms of small models is mostly a mirage: even at 20x cheaper, a small model likely spends 40-50x more tokens to solve a problem than the bigger one. A quoted reply adds a real test: Haiku 5.5 at xhigh ended up more expensive than Opus at medium — it just kept generating while left unattended.
More from Models
- Dev debate: is ColBERT-style late interaction still a cross-encoder as rerankers fade? — CShorten30 · 2026-10-08
- Claude Projects quietly adds scheduled tasks for automated recurring runs — ColleenMBrady · 2026-10-08
- Leak: X preps all-in-one subscription bundling X, Grok and Cursor in one usage pool — nima_owji · 2026-10-08
- 4 models, one two-file bug: 3/4 passed, 10x cost spread, and the cheapest run was the failure — lulzxdxdxd · 2026-10-08
- Google Moves Gemini Flash and Pro to Paid Plans, Free Tier Keeps Flash-Lite — Robert__Sinclair · 2026-10-08
- Apollo Research: Final-Checkpoint Evals Can't Catch Misalignment That Emerges Early — dl_weekly · 2026-10-08