Xiaomi MiMo-V2.6-Pro fixes 22.7% of 105 real bugs for just $0.86 in blind test

PawelHuryn · x · 2026-09-22

Pawel Huryn benchmarked Xiaomi's newly released MiMo-V2.6-Pro on real work: 2 repos, 105 actual bugs, blind-judged scoring.

Takeaway: MiMo-V2.6-Pro and DeepSeek V4.1 Flash form a strong Pareto frontier — near-frontier scores at roughly an order of magnitude lower cost. The author argues these models would remain viable daily drivers even if subsidized subscriptions disappeared. All scores are medians of 3+ runs except Luna (n=2).

Related event: Bug Hunt Bench Tests Models on 105 Real Bugs, Cost Gap Nears 200x(3 posts)→

Original post →

More from Models

Models channel →