Kimi K3 beats GLM-5.2 in exploit tests but still fails end-to-end attacks
kevinsxu · x · 2026-07-28
The post quotes the Kimi K3 paper’s cyber evaluation results and notes that a joint UK AI Security Institute / NIST CAISI assessment reached similar conclusions.
- Kimi K3 outperforms GLM-5.2 on exploit development.
- It scores 32% vs. 24% on ExploitBench.
- On a simulated enterprise network task that takes a human expert about 20 hours, it needs 17 steps vs. 11 steps.
- But it still trails frontier cyber-capable models on end-to-end exploit completion, achieving arbitrary code execution on 0 of 41 tasks.
The punchline is that China and the US institutions agree on the ranking, at least on this slice of cyber capability.
Related event: Kimi K3 Lags in Red Teaming Despite Exploit Gains(2 posts)→
More from Models
- Opus 5 Review: Technically Brilliant but Stiff, Tailored for Sub-Agents Not Solo Chat — MicahBerkley · 2026-07-28
- Google AI reaches a bizarre conclusion in a screenshot that Reddit can't ignore — FrankieMint · 2026-07-28
- AMD 6800H APU benchmark finds Qwen 3.6 MoE far faster than 31B Q8_0 — tabletuser_blogspot · 2026-07-28
- Kimi K3 now has a public viewer with a breakdown of 896 experts — uncommoncrawl · 2026-07-28
- Kimi K3 Architecture Visualizer Launches with Deep Dive into 896 Experts — Course_Latter · 2026-07-28
- Big models are “wiser” while more reasoning makes them more “diligent” — breath_mirror · 2026-07-28