UK evaluators say Kimi K3 lags U.S. frontier models on cyber capability tests
BlackHC · x · 2026-07-24
UK AISI and CAISI report that Kimi K3 trails frontier U.S. models on cyber capability evaluations.
- On ExploitBench, Mythos Preview reached arbitrary code execution in 18/41 tasks, while Kimi K3 never did.
- On “The Last Ones,” Mythos Preview solved the benchmark on 3/10 tries with an average score of 22/32; Kimi K3 solved it once and averaged 17/32.
- The evaluators also say Kimi K3’s safeguards did not stop it from attempting exploit development or offensive cyber operations during testing.
More from Models
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11