Ramp opens its LLM router after cutting AI costs by 30%+
iamrobotbear · x · 2026-07-21
Ramp says its internal LLM router is now being opened up as Ramp Router: one OpenAI-compatible endpoint that sends each request to the cheapest model that still clears the quality bar.
- The system tests new models on real work before routing traffic.
- It also decides when to cache, batch, or escalate to a stronger model.
- Ramp says the approach has already cut costs by 30%+ while keeping output the same or better.
- The alpha group is reportedly running at 2.75T tokens/month with 99.99% uptime.
The post frames model routing as the practical answer to fast-moving prices and capabilities across GPT, Claude, Gemini, Grok, Qwen, DeepSeek, Kimi, and GLM.
Related event: Ramp Opens Internal Multi-Model Router, Claims 30% LLM Cost Reduction(11 posts)→
More from Infra
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11