GLM 5.3 Flash hits RunInfra day one: 254 tok/s, 1M context, $0.10/1M input
alejandroll10 · x · 2026-08-27
RunInfra has launched GLM 5.3 Flash (vendor-native FP8 from zai-org) on day zero, competing on speed, price and caching:
- 254.1 output tok/s, 703ms to first token, 1M-token context window
- Pricing: $0.10/1M input, $0.40/1M output, cached input at $0.01/1M with a reported 86–92% cache hit rate, making repeated prefixes nearly free
- OpenAI-compatible API — just change the base URL; tool calling, JSON mode and streaming supported
- Zero data retention: prompts are never stored or used for training
More from Infra
- Prediction: Consumer Desktops Will Soon Run k3-Quality Models — AaronBergman18 · 2026-08-27
- Anthropic Locks in 460MW Compute for $45B, Revealing GPU Economics — zephyr_z9 · 2026-08-27
- Kioxia Plans $6.27B Investment for Third Fab in Iwate — zephyr_z9 · 2026-08-27
- DeepSeek-V4-Flash hits 51.5 tok/s on M3 Ultra — antirez · 2026-08-27
- Colibrì Engine Update: Runs 2.8T Param Models, Boosts Speed via Expert Caching — solyarisoftware · 2026-08-27
- Jensen Huang: Data Center Investment Payback Period Under One Year — zephyr_z9 · 2026-08-27