Smart LLM routing cuts costs 69% on 120 tasks while keeping 99.2% success rate
shensi · x · 2026-09-09
Merge ran their Gateway evals using only first-party models from Anthropic, OpenAI, and Google to test the assumption that model routing saves money only by swapping in cheaper open-source models. With smart routing, the same 120 tasks cost 69% less, returned faster every time, and still hit a 99.2% success rate.
Key points:
- Even routing exclusively among closed frontier models yields major cost and latency savings.
- Delivered via Merge Gateway: one API to access multiple LLMs with built-in routing, cost management, and security.
- SDK examples show routing targets like openai/gpt-5.6-luna, anthropic/claude-haiku-4-5, and google/gemini-3.1-flash-lite.
It's a vendor-published eval, but the concrete numbers make it useful reference for production LLM gateway decisions.
More from coding & agent
- Best free open-source AI tools of the year: CloakBrowser, curated lists, efficient coding agents — 0xsachi · 2026-09-10
- Harness, a physical device to manage all your coding agents, ships Friday — dee_hw · 2026-09-09
- Developer builds a computer vision tennis coach tracking ball speed and stroke form — measure_plan · 2026-09-09
- Replacing a real HVAC company's entire SaaS stack with one vertically integrated agent — _AustinCalvert_ · 2026-09-09
- Solo devs: what do you ship when the model confidently misreads user data? — FamiliarSlide7685 · 2026-09-09
- Breakdown: how many tokens a monthly coding agent subscription actually buys — zainhas · 2026-09-09