UK evaluators say Kimi K3 lags U.S. frontier models on cyber capability tests
BlackHC · x · 2026-07-24
UK AISI and CAISI report that Kimi K3 trails frontier U.S. models on cyber capability evaluations.
- On ExploitBench, Mythos Preview reached arbitrary code execution in 18/41 tasks, while Kimi K3 never did.
- On “The Last Ones,” Mythos Preview solved the benchmark on 3/10 tries with an average score of 22/32; Kimi K3 solved it once and averaged 17/32.
- The evaluators also say Kimi K3’s safeguards did not stop it from attempting exploit development or offensive cyber operations during testing.
More from Models
- Kimi-k3 lands 1 point behind fable-5 on EQBench and tops writing tests — karminski3 · 2026-07-24
- Fable 5 reaches 67.3% on GameDevBench, trailing only GPT-5.6 Sol — scaling01 · 2026-07-24
- Grok 4.5 scores 85.7% on ARC-AGI-1 but just 0.3% on ARC-AGI-3 — scaling01 · 2026-07-24
- GPT-5.6 deletes an Act 1 boss in two turns on a live stream — Jsevillamol · 2026-07-24
- Kimi K3 lands on Together AI at launch for coding and agent workloads — togethercompute · 2026-07-24
- User says ChatGPT 5.6 Pro helped disprove a 22-year-old graph theory conjecture — basedjensen · 2026-07-24