UK AISI/CAISI finds Kimi K3 lags frontier cyber models in attack tests
ShakeelHashim · x · 2026-07-24
Kimi K3 trails frontier cyber models in UK AISI/CAISI tests
A joint UK AISI / CAISI evaluation found that Kimi K3 performs significantly below the latest frontier cyber-capable models on preliminary cyber tests.
- On exploit-development tasks, Kimi K3 ranked well below leading frontier models.
- On the simulated corporate-network attack benchmark “The Last Ones”, it averaged step 17 of 32, while the most cyber-capable U.S. models averaged 28.5 steps.
- The report says Kimi K3 still outperformed GLM-5.2 on the same evaluations.
- Its safeguards did not prevent attempts at cyber exploit development or offensive cyber operations during the tests.
Related event: UK and US Safety Institutes Find Kimi K3 Lags in Cybersecurity Capabilities(4 posts)→
More from Models
- Rumor says Anthropic’s Opus 5 has slipped to tomorrow — mark_k · 2026-07-24
- Epoch AI Live-Streams GPT-5.6 Playing Slay the Spire — Jsevillamol · 2026-07-24
- NVIDIA introduces NVFP4 for faster LLM inference with less GPU memory — NVIDIA Developer · 2026-07-24
- ChatGPT Stuck for 10 Minutes: Long Context Threads Hit Stability Wall — billyjhowell · 2026-07-24
- Dev Projects Blocked by Safety Filters: A Push for Open Weights — VoidStateKate · 2026-07-24
- Reddit side-by-side test says GPT-5.6 SOL beats KIMI 3 on Chinese ink-wash animation — notNIHAL · 2026-07-24