GLM 5.3 Flash kernels rewritten on RunInfra: 670 tok/s, 99.7% cache hit, AMD support
ycombinator · x · 2026-10-01
RunInfra spent September rewriting the inference kernels behind GLM 5.3 Flash and shipped the major release, hitting 670 tok/s on Vercel AI Gateway.
- Pricing: $0.11/1M input, $0.45/1M output, $0.03/1M cached tokens
- Context: 1M tokens, FP8 vendor-native build
- Hardware: now runs on AMD GPUs — same model, same API, more capacity
- Cache: 99.7% hit rate over the last 24 hours, with per-hit/miss logs in the dashboard
- Compatibility: OpenAI-compatible chat completions plus Anthropic-compatible /v1/messages, with image input, tool calling, JSON mode and streaming
- Privacy: zero data retention, never used for training
More from Infra
- Magnitude (YC S25) Open-Sources Inference Engine with Up to 2x Faster Decode Than llama.cpp — petrusenko_max · 2026-10-01
- Chinese labs reportedly buying expert training data from US data firms like Mercor and SurgeAI — BarnacleBasic1988 · 2026-10-01
- Andrew Chen: The $20K workstation era is back — RTX PRO 6000 and Mac Studio — andrewchen · 2026-10-01
- PyTorch Foundation's Mark Collier: open source is the coordination layer for frontier AI — PyTorch · 2026-10-01
- Qwen 27B Now Runs on AMD NPUs via FastFlowLM, at a Slow 1 tps — TuskNaPrezydenta2020 · 2026-10-01
- Cognition Becomes First CoreWeave Vera Rubin NVL72 Customer, Sees 4.8X SWE-2 Throughput Boost — altryne · 2026-10-01