Ramp Launches Router: LLM Gateway Cuts Inference Costs by 40%
SethGRosenberg · x · 2026-08-28
Ramp released 'Router', an LLM gateway designed to reduce inference costs by 40% on average by routing every request to the lowest-cost model that meets performance requirements. It consolidates access to multiple models (closed and open-source) behind a single API endpoint, supports automatic routing, and continuously rolls out new cost-saving strategies. Customer Delphi reported a 92% reduction in model costs after routing billions of tokens. The tool includes a CLI and offers free routing through 2026.
More from Infra
- AI automated research finds numerical bug in vLLM/SGLang backend — josh_tobin_ · 2026-08-28
- Custom llama.cpp branch speeds up Metal Qwen3.8-Flash-Next inference, adds n-gram SSD offload — tarruda · 2026-08-28
- Hot Chips 2026: inference chips enter an "era of ferment" with divergent bets — BenBajarin · 2026-08-28
- Testing Muon Optimizer: Smoother Gradients and Stable Residual Maxima — stochasticchasm · 2026-08-28
- Analyst predicts CXL standard commercialization to break the memory wall — BenBajarin · 2026-08-28
- A Deep Dive Into China's HBM/DRAM Situation: Memory as the New Bottleneck — demian_ai · 2026-08-28