Xiaomi MiMo-V2.6-Pro fixes 22.7% of 105 real bugs for just $0.86 in blind test
PawelHuryn · x · 2026-09-22
Pawel Huryn benchmarked Xiaomi's newly released MiMo-V2.6-Pro on real work: 2 repos, 105 actual bugs, blind-judged scoring.
- Muse Spark 1.3 (max): 32.2 for $18.11
- GPT-5.6 Luna (max): 31.5 for $2.82
- Grok 4.7 (xhigh): 28.8 for $22.89
- MiMo-V2.6-Pro (default): 22.7 for $0.86
- DeepSeek V4.1 Flash (max): 21.7 for $0.78
- Gemini 3.8 Flash (high): 18 for $11.03
Takeaway: MiMo-V2.6-Pro and DeepSeek V4.1 Flash form a strong Pareto frontier — near-frontier scores at roughly an order of magnitude lower cost. The author argues these models would remain viable daily drivers even if subsidized subscriptions disappeared. All scores are medians of 3+ runs except Luna (n=2).
Related event: Bug Hunt Bench Tests Models on 105 Real Bugs, Cost Gap Nears 200x(3 posts)→
More from Models
- Sentdex corrects himself: 9GB memory serves the default 4B Qwen model — Sentdex · 2026-09-22
- Rich RL Report Ships With 9B Distilled Model and 7,000 Open RL Environments — tokenbender · 2026-09-22
- Anthropic's Real-World AI Intrusion Data on APT29 Deserves More Attention Than Doom Debates — jcran · 2026-09-22
- Dev questions whether Prime Intellect's livestreamed RL run was a reenactment — giffmana · 2026-09-22
- Thermofluids professor: OpenAI's Clay problem solution is not physically reproducible — GaryMarcus · 2026-09-22
- A local 27B Qwen agent running on 24GB DDR4 started acting human after getting long-term memory — Composer_Fearless · 2026-09-22