CAISI says Kimi K3 trails leading U.S. frontier models in cyber capability
burny_tech · x · 2026-07-25
CAISI’s latest blog post evaluates Kimi K3 on cyber capability and says it performs significantly below the leading U.S. frontier models.
The attached chart plots a cyber-capability Elo trend for U.S. vs. PRC models, with Kimi K3 sitting below other recent Chinese releases such as GLM 5.2 and DeepSeek V4 Pro.
Related event: UK AISI/CAISI Evaluate Kimi K3: Trails US Frontier Models in Cyber(11 posts)→
More from Models
- ‘The only eval I pay attention to’: Opus 5 and a RunescapeBench joke — Flomerboy · 2026-07-25
- GPT 5.6 Becomes More Casual in Chat, Opus 5 Can Feel Testy — dreamwieber · 2026-07-25
- Claude Opus 5 Hits 70.6% on OSWorld v2, Dev Offers Bounty for Harder Evals — EricBuess · 2026-07-25
- Miles Brundage says Opus 5 copies joke setups, then offers to “be original” — Miles_Brundage · 2026-07-25
- Kimi K3.0 abliterated GGUF starts trending on Hugging Face — audnai · 2026-07-25
- Early reaction to Opus 5: colder, less pushback than the 4.7–4.8 line — teortaxesTex · 2026-07-25