Smart routing across frontier labs only: 69% cost cut, 99.2% success, lower latency
shensi · x · 2026-09-10
Answering skeptics who assume model routing saves money only by swapping in cheaper open-source models, Merge API re-ran its gateway eval restricted to frontier-lab models from Anthropic, OpenAI and Google. Across the same 120 tasks, smart routing cut costs 69%, returned results faster every time, and still hit a 99.2% success rate — showing big savings exist even within frontier models.
Related event: Smart Routing Cuts Costs 69% Using Only Frontier Models(2 posts)→
More from Infra
- NASA chief backs orbital AI compute as SpaceX targets first space data center in 2027 — rohanpaul_ai · 2026-09-10
- NVIDIA ships open-source PAIR: turns idle home PCs into a local AI cluster, ~51% faster in demo — solyarisoftware · 2026-09-10
- Massachusetts moves to require community agreements before data center permits — Polymarket · 2026-09-10
- At Scale, KV Cache Becomes a Storage System: How LLMs Serve GBs of Cached State — blaizedsouza · 2026-09-10
- Don't let FOMO win: you can learn more about local LLMs with a tiny model than a 5090 — sn2006gy · 2026-09-10
- Stealth Chip Startup Kepler Emerges With EUV-Free 3D-Stacked HBM, $468M Raised — pstAsiatech · 2026-09-10