UK AISI/CAISI finds Kimi K3 lags frontier cyber models in attack tests
ShakeelHashim · x · 2026-07-24
Kimi K3 trails frontier cyber models in UK AISI/CAISI tests
A joint UK AISI / CAISI evaluation found that Kimi K3 performs significantly below the latest frontier cyber-capable models on preliminary cyber tests.
- On exploit-development tasks, Kimi K3 ranked well below leading frontier models.
- On the simulated corporate-network attack benchmark “The Last Ones”, it averaged step 17 of 32, while the most cyber-capable U.S. models averaged 28.5 steps.
- The report says Kimi K3 still outperformed GLM-5.2 on the same evaluations.
- Its safeguards did not prevent attempts at cyber exploit development or offensive cyber operations during the tests.
More from Models
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11