Ramp opens its LLM router after cutting AI costs by 30%+
iamrobotbear · x · 2026-07-21
Ramp says its internal LLM router is now being opened up as Ramp Router: one OpenAI-compatible endpoint that sends each request to the cheapest model that still clears the quality bar.
- The system tests new models on real work before routing traffic.
- It also decides when to cache, batch, or escalate to a stronger model.
- Ramp says the approach has already cut costs by 30%+ while keeping output the same or better.
- The alpha group is reportedly running at 2.75T tokens/month with 99.99% uptime.
The post frames model routing as the practical answer to fast-moving prices and capabilities across GPT, Claude, Gemini, Grok, Qwen, DeepSeek, Kimi, and GLM.
Related event: Ramp Opens Internal Multi-Model Router Ramp Router to the Public(11 posts)→
More from Infra
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22
- An energy model says the U.S. could run short of natural gas starting in 2028 — churchkey · 2026-07-22
- Arbitrum fee simulation shows higher gas capacity but lower L2 revenue under ArbOS61 — tomwanhh · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22