Kimi K3 Cybersecurity Eval Sparks Debate: Scores 32.2%
Recent cybersecurity evaluations of Kimi K3 have sparked widespread discussion. According to ExploitBench data, the model achieved a total score of only 32.2% with zero ACEs (Autonomous Complete Exploits) across 41 samples, lagging significantly behind top US models that average 76.2%, leading to speculation about whether Moonshot submitted an optimized version for testing.
Confirmed
In the ExploitBench evaluation, Kimi K3 scored 32% (or 32.2%), with 0 out of 41 samples achieving ACE. According to research cited by SCMP, the average score for the best-performing US models is 76.2%.
Unconfirmed
There is debate over Kimi K3's exact ranking in overall cybersecurity capabilities. Some comparative charts (provided by @zainhas) show Kimi K3 leading among available models, surpassing GLM 5.2 and DeepSeek V4 Pro. However, charts from @petrusenkomax indicate that on the cyber range metric, which is closer to real-world attack capabilities, Kimi K3 still trails Mythos Preview. Additionally, the claim that Moonshot did not provide its optimal version to evaluators remains an external speculation.
Why it matters
This evaluation exposes the gap between top Chinese and US LLMs in foundational cybersecurity attack and defense capabilities. However, @dyn notes that while calling this score a direct measure of "cyber capability" is somewhat exaggerated, and the model's overall cyber prowess is weak, Kimi K3 can already execute complete long-chain tasks, marking its engineering feasibility in complex security operations.
2026-07-24 ~ 2026-07-25 · 5 related posts
- Episode 1: Kimi K3 Cybersecurity Eval Sparks Debate: Scores 32.2%(2026-07-24, 5 posts)
- Episode 2: Kimi K3 Security Eval: High Exploit, Low Guardrails(2026-07-26, 4 posts)
Primary sources
- [source] Kimi K3 scores 32% on ExploitBench and reaches 0 of 41 ACE cases — xeophon · 2026-07-24
- Kimi K3’s cyber eval still looks weak, but it can finish a full long-horizon chain — dyn___ · 2026-07-24
- Kimi K3 trails Mythos on cyber-range tasks in a benchmark chart shared online — petrusenko_max · 2026-07-25
- [source] Study says Kimi K3 scores 32.2% on cyberattack ability vs 76.2% for U.S. models — pstAsiatech · 2026-07-25
- [source] Chart puts Kimi K3 at the top of publicly usable cyber models — zainhas · 2026-07-25