INT21 claims 20 AI-generated inference engines in 2 weeks, MiMo hits 1,308 tok/s

bingxu_ · x · 2026-09-29

Bing Xu shared INT21's inference benchmark, arguing human-crafted AI software is becoming history. INT21 positions ultra-fast inference as the bottleneck of the agentic era: a coding agent makes dozens of sequential calls per fix. Two engineers directed generation of 20 engines in 2 weeks across 7 model categories (autoregressive/diffusion LM, ASR, TTS, OCR, image, video). On a 15-model subset vs SGLang and vLLM, INT21 leads decode rate on all 6 text models tested: MiMo at 1,308 tok/s vs 540 for tuned SGLang and 1,011 for vLLM.

Original post →

More from coding & agent

coding & agent channel →