GLM 5.3 Flash kernels rewritten on RunInfra: 670 tok/s, 99.7% cache hit, AMD support

ycombinator · x · 2026-10-01

RunInfra spent September rewriting the inference kernels behind GLM 5.3 Flash and shipped the major release, hitting 670 tok/s on Vercel AI Gateway.

Original post →

More from Infra

Infra channel →