Kimi K3 Cybersecurity Eval Sparks Debate: Scores 32.2%

Recent cybersecurity evaluations of Kimi K3 have sparked widespread discussion. According to ExploitBench data, the model achieved a total score of only 32.2% with zero ACEs (Autonomous Complete Exploits) across 41 samples, lagging significantly behind top US models that average 76.2%, leading to speculation about whether Moonshot submitted an optimized version for testing.

Confirmed

In the ExploitBench evaluation, Kimi K3 scored 32% (or 32.2%), with 0 out of 41 samples achieving ACE. According to research cited by SCMP, the average score for the best-performing US models is 76.2%.

Unconfirmed

There is debate over Kimi K3's exact ranking in overall cybersecurity capabilities. Some comparative charts (provided by @zainhas) show Kimi K3 leading among available models, surpassing GLM 5.2 and DeepSeek V4 Pro. However, charts from @petrusenkomax indicate that on the cyber range metric, which is closer to real-world attack capabilities, Kimi K3 still trails Mythos Preview. Additionally, the claim that Moonshot did not provide its optimal version to evaluators remains an external speculation.

Why it matters

This evaluation exposes the gap between top Chinese and US LLMs in foundational cybersecurity attack and defense capabilities. However, @dyn notes that while calling this score a direct measure of "cyber capability" is somewhat exaggerated, and the model's overall cyber prowess is weak, Kimi K3 can already execute complete long-chain tasks, marking its engineering feasibility in complex security operations.

2026-07-24 ~ 2026-07-25 · 5 related posts

Full story(2 episodes)→

Primary sources