INT21 claims 20 AI-generated inference engines in 2 weeks, MiMo hits 1,308 tok/s
bingxu_ · x · 2026-09-29
Bing Xu shared INT21's inference benchmark, arguing human-crafted AI software is becoming history. INT21 positions ultra-fast inference as the bottleneck of the agentic era: a coding agent makes dozens of sequential calls per fix. Two engineers directed generation of 20 engines in 2 weeks across 7 model categories (autoregressive/diffusion LM, ASR, TTS, OCR, image, video). On a 15-model subset vs SGLang and vLLM, INT21 leads decode rate on all 6 text models tested: MiMo at 1,308 tok/s vs 540 for tuned SGLang and 1,011 for vLLM.
More from coding & agent
- Agents can exploit nondeterminism to modify sandboxes beyond what traces reveal — lbeurerkellner · 2026-09-29
- PhD defense slides with dead JS deps resurrected by Claude in one shot — Ben_Reinhardt · 2026-09-29
- Developer coins 'stochastic productivity': agentic coding output swings wildly day to day — carsonfarmer · 2026-09-29
- Dev open-sources 8 Godot agent skills distilled from building a game with AI in Claude Code/Codex — AIandDesign · 2026-09-29
- Cekura speech benchmark: 2,214 live calls, GPT Realtime 2.1 most reliable, Phonic v1 fastest at 1.59s — graceisford · 2026-09-29
- Before building an AI agent, ask the human doing the work what's not in the docs — alex_verem · 2026-09-29