LLM routing saved 33.2% vs premium models in 640-request pilot, but a fixed mid-priced model beat it
smakosh · reddit · 2026-09-26
LLM Gateway self-benchmarked its Jev-based routing against fixed-model baselines across 640 measured requests (80 prompts, four arms, two output limits). Routing cost 33.2% less than a premium baseline but passed 75/79 tasks vs 79/79 after adjusted scoring, and a fixed mid-priced model cost 72.9% less than routing while passing 76/79. The cost confidence interval was wide (4.3% higher to 66.1% lower), and the pilot excluded long conversations, tool use, and agentic coding. Takeaway: routing trims premium bills, but a cheaper fixed model remains a strong alternative. Full prompts, costs, and reproducible scripts are published.
More from Infra
- Why rent servers when agents can run your terminal? — StewartalsopIII · 2026-09-26
- MIT paper: scaling law expiring as cost doubles 6x per width jump — DavidLinthicum · 2026-09-26
- Investor thesis: GOES steel and transformer makers may outshine rare earths — basedjensen · 2026-09-26
- Polymorf pushes OMLX inference from 150tps to nearly 200tps within 24 hours of launch — HankYeomans · 2026-09-26
- FlashLoop exploits cross-loop redundancy to speed up Looped Transformers by 1.65x — KyeGomezB · 2026-09-26
- Why this builder quit server racks: fried motherboards and a ~$1,500 housing bill — TheZachMueller · 2026-09-26