Ramp launches Router LLM gateway to cut inference costs by 40%
round · x · 2026-08-22
Ramp has officially launched Router, an LLM gateway designed to reduce inference costs by intelligently routing requests to the most cost-effective model that meets performance needs. The company claims an average reduction of 40% in AI costs.
Key Features:
- Unified Access: Access all models through a single API key and bill.
- Smart Routing: Automatically switches models to save costs with an enabled auto mode.
- Cost Optimization: A case study highlights a customer reducing model costs by 92% while running billions of tokens through Router.
- Compatibility: Supports both closed and open-source models, US-hosted with Zero Data Retention (ZDR) options.
The product targets both CTOs for optimal model selection and CFOs for lower spending.
More from Infra
- Alibaba & ByteDance paper: Model inference is no longer the main bottleneck for AI agents — rohanpaul_ai · 2026-08-22
- Opinion: GPUs are massively underpriced given their intelligence value — Technical_Ad_6106 · 2026-08-22
- Scratch-built engine beats vendor runtime for ternary 8B models on free ARM cores — Annual_Manner_5901 · 2026-08-22
- Nvidia's New Vera CPU Tested: Big Bandwidth & FP8 Boost Agent Execution — pzakin · 2026-08-22
- Green Compute Launches Biogas-Powered GPU Cluster on Bittensor with 4090/5090 Rentals — markjeffrey · 2026-08-22
- SilkStack v2.2: local semantic search with custom WebLLM embedding model — skk80 · 2026-08-22