Kimi K3 Evaluations Show Polarized Results and Harness Sensitivity

Moonshot's Kimi K3 model has recently sparked widespread evaluation and discussion within the AI community. The model's performance shows significant polarization across different benchmarks, highlighting the complexity of current AI evaluation systems and drawing heavy attention to the impact of testing harnesses.

Key Details and Performance Polarization

Kimi K3's scores exhibit a clear two-tier polarization. On one hand, data provided by @ChrisUniverse shows K3 ranking 1st in `Arena Frontend Code` with a score of 1679. On the other hand, it scored only 39% in the `FrontierMath` Tier 4 benchmark (@scaling01), which is 7% lower than the best US models from 7 months ago. Furthermore, K3 ranked at the bottom in real-world code repair tests (@ChrisUniverse), contradicting claims of its strong coding capabilities.

The Impact of Testing Harnesses

Addressing the controversy over its practical performance, several authors point out that the testing harness massively influences K3's results. @victormustar emphasized that for the exact same task, switching evaluation frameworks could change K3's performance from "terrible" to near Fable-level, yielding stunning results in Boeing 747-related tasks. In a specific bug-fixing comparison (reposted by @ssh4net), both Kimi K3 and Fable 5 successfully fixed real bugs, but Fable 5 was more efficient, taking 3.5 minutes and requiring 18 tool calls.

Interpretations and Perspectives

Regarding the low math evaluation scores, @teortaxesTex refuted the notion that "Chinese models are losing momentum." They argued that frontier mathematics is historically an area where Chinese models lag, and this low score primarily dragged down the model's ECI estimate. However, they predicted that Chinese models typically catch up rapidly in subsequent versions.

2026-07-18 ~ 2026-07-20 · 5 related posts

Full story(20 episodes)→