Together serves 23%-30% of all OpenRouter traffic for GLM 5.3 models

zhyncs42 · x · 2026-09-12

Together AI says it is serving GLM 5.3 and GLM 5.3 Flash at top-decile TPS, latency and cache rates, handling 23% and 30% of all OpenRouter traffic for the two models respectively — with OpenRouter only a fraction of its total API volume. The team stresses scaled agentic inference requires optimizing all dimensions with reliability, not maxing a single metric. Per OpenRouter, GLM 5.3 is Z.ai's large reasoning model for software engineering and long-horizon agent tasks with 1M-token context, priced at $0.8727/$3.36 per 1M tokens.

Original post →

More from Infra

Infra channel →