OpenRouter Introduces Flex Tier
strickvl · x · 2026-07-11
OpenRouter has introduced a new flex service tier for models.
It acts somewhat like a discounted version of /batch: eligible models receive a 50% price discount without the need to handle a full asynchronous queue like the batch mode.
The trade-off is that the client may need to handle rate limiting more proactively, and latency will be slower and more volatile. The post mentions that for eligible Gemini models, this tier effectively cuts the price in half.
More from Infra
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22
- Report says Nvidia could build 1,000 Vera Rubin racks a day, implying $630B quarterly at system level — GavinSBaker · 2026-07-22
- oMLX 0.5.2 adds Mac menu-bar stats, low-bit decode kernels, and faster downloads — awnihannun · 2026-07-22