AI-generated inference engine hits 1300 tok/s on Xiaomi's 1T model, beating SGLang and vLLM

bingxu_ · x · 2026-09-30

INT21's two-engineer team spent two weeks generating 20 inference engines and benchmarked 15 of them against SGLang and vLLM. Standouts: Xiaomi MiMo V2.6 Pro at 1308 tok/s decode on 8×B300 (vs 540/1011), and DeepSeek V4.1 Flash at 1844 tok/s on 8×B200 — a 2.22×/2.33× speedup. The report, titled "An AlphaGo Moment for Inference?", publishes all measurements and notes limits like higher first-token latency on some workloads.

Related event: Two-Person Team Uses AI to Generate 20 Inference Engines in Two Weeks, Hitting 1300 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →