OpenRouter routing distorts inference pricing: GLM 5.3 at $0.08/M input vs $5/M output
soumithchintala · x · 2026-10-05
cHHillee flags an interesting distortion in inference pricing apparently driven by OpenRouter's routing: ridiculously cheap prefill and very expensive decode.
- GLM 5.3's highest-volume provider, inference.net, charges $0.08/M input tokens but $5.00/M output — a 60x+ gap
- This means input-heavy workflows (long context, caching-style usage) are dirt cheap while long generations get expensive fast
- Soumith Chintala's amplification suggests the observation is resonating among engineers
Directly useful for provider selection and cost estimation.
More from Infra
- A 1 GW fusion reactor could breed ~2 tonnes of gold a year from mercury, half its revenue — anselm · 2026-10-05
- New Substack essay draws parallels between the cannabis industry and data centers — ctjlewis · 2026-10-05
- Ibiden: ABF substrate supply can't keep up as AI silicon area balloons — zephyr_z9 · 2026-10-05
- Only 2% of US electricians are DC-certified, says Ben Horowitz — AI infra bottleneck — rohanpaul_ai · 2026-10-05
- United Airlines to equip 880+ aircraft with Starlink by end of 2026 — elonmusk · 2026-10-05
- Nutanix Buys Ryax to Squeeze More Work Out of Enterprises' Idle GPUs — shashib · 2026-10-05