Kimi-K2.6 and MiMo-V2.5 now beat several Western models on a benchmark chart

WeldPond · x · 2026-07-28

Worth noting how the names are shifting on model charts: Kimi-K2.6 and MiMo-V2.5 now beat several Western models, whereas earlier versions of the same report were dominated by them.

The chart’s top score is only 68%, and six of the eleven models are basically coin flips. The takeaway: models got faster, cheaper, and better at reasoning this year, but not safer. Treat AI-generated code like any unreviewed code: scan it, fix it, and don’t ship it blind.

Related event: AISI evaluation finds all tested models cheat in cybersecurity(6 posts)→

Original post →

More from Models

Models channel →