Together serves 23%-30% of all OpenRouter traffic for GLM 5.3 models
zhyncs42 · x · 2026-09-12
Together AI says it is serving GLM 5.3 and GLM 5.3 Flash at top-decile TPS, latency and cache rates, handling 23% and 30% of all OpenRouter traffic for the two models respectively — with OpenRouter only a fraction of its total API volume. The team stresses scaled agentic inference requires optimizing all dimensions with reliability, not maxing a single metric. Per OpenRouter, GLM 5.3 is Z.ai's large reasoning model for software engineering and long-horizon agent tasks with 1M-token context, priced at $0.8727/$3.36 per 1M tokens.
More from Infra
- Zilliz CTO: Agent memory is a long-lived systems problem, not an index feature — J_Luan_ · 2026-09-12
- Draw Things update adds MiniMax H3 with LoRA/TeaCache and Krea 2 model imports — antirez · 2026-09-12
- Wafer launches 'most comprehensive' AI performance engineering repo, starting with Transformer inference deep-dive — ycombinator · 2026-09-12
- The Global Race for Cheap Power: Where AI Data Centers Should Actually Go — pravchaw · 2026-09-12
- VCs float 'hardware revenue derivative': fund compute costs via revenue share, not equity — ns123abc · 2026-09-12
- Relace hits 1T tokens/day on OpenRouter, serving 37% of DeepSeek v4 Flash traffic — stuffyokodraws · 2026-09-12