Relace hits 1T tokens/day on OpenRouter, serving 37% of DeepSeek v4 Flash traffic
stuffyokodraws · x · 2026-09-12
Inference provider Relace announced it now serves over 1 trillion tokens per day via OpenRouter, breaking down its share by model: 37% of DeepSeek v4 Flash traffic, 25% of GLM 5.3 Flash, and just 1.6% of Kimi K3.
The numbers offer a rare glimpse into real-world API traffic distribution across Chinese open-weight models and the scale of third-party inference serving.
More from Infra
- Wafer launches 'most comprehensive' AI performance engineering repo, starting with Transformer inference deep-dive — ycombinator · 2026-09-12
- Together serves 23%-30% of all OpenRouter traffic for GLM 5.3 models — zhyncs42 · 2026-09-12
- The Global Race for Cheap Power: Where AI Data Centers Should Actually Go — pravchaw · 2026-09-12
- VCs float 'hardware revenue derivative': fund compute costs via revenue share, not equity — ns123abc · 2026-09-12
- Yutori's Navigator n2 runs browser agents at $1.46 per task on OSWorld 2.0 vs $13-$40+ for frontier models — DhruvBatra_ · 2026-09-12
- Auto-derived FlashAttention with SMEM and tensor core assignment shown off — vtabbott_ · 2026-09-12