AI-generated inference engine hits 1300 tok/s on Xiaomi's 1T model, beating SGLang and vLLM
bingxu_ · x · 2026-09-30
INT21's two-engineer team spent two weeks generating 20 inference engines and benchmarked 15 of them against SGLang and vLLM. Standouts: Xiaomi MiMo V2.6 Pro at 1308 tok/s decode on 8×B300 (vs 540/1011), and DeepSeek V4.1 Flash at 1844 tok/s on 8×B200 — a 2.22×/2.33× speedup. The report, titled "An AlphaGo Moment for Inference?", publishes all measurements and notes limits like higher first-token latency on some workloads.
More from Infra
- SGLang turns Qwen3.8-27B into a decision model that beats Pokémon FireRed at sub-100ms — zhaoran_wang · 2026-09-30
- Agentic AI turns CPUs into the overlooked bottleneck as CPU:GPU ratios shift upward — AccBalanced · 2026-09-30
- AT&T CEO says SpaceX's phone strategy is not viable: "Satellite won't beat fiber" — RachelVT42 · 2026-09-30
- ModelScope CLI quietly moved to the modelscope-hub package — PSA to save you 30 minutes — pilkyton · 2026-09-30
- Rowmax-H15: approximate softmax in attention for 25.8% faster B200 inference with minimal quality loss — illinois · 2026-09-30
- AWS Bedrock Brings Claude Opus 5, Sonnet 5, Haiku 4.5 to In-Country Inference in India — AWS ML Blog · 2026-09-30