UK/US Agencies Evaluate Kimi K3: Lags Behind Frontier Models in Cyber Capabilities
A joint preliminary assessment by UK AISI and US CAISI evaluated the cyber capabilities of Moonshot's Kimi K3. The findings reveal that Kimi K3 significantly lags behind US frontier models in cyber-related tasks, though it leads among open-weight models.
Confirmed
According to information shared by the agencies and related sources, the tests found that Kimi K3's performance is "significantly lower" than US frontier models. In specific cyber capability evaluations, it fell short of the latest generation of frontier models in offensive cyber tasks. For instance, in the ExploitBench test, Kimi K3 stopped at an average of 17 steps; by comparison, the frontier model Mythos Preview achieved Arbitrary Code Execution (ACE) in 18 out of 41 tasks. Additionally, charts from the Commerce / CAISI-NIST report indicate that the comparison included other models like GLM-5.2. The evaluation was noted as a fast-tracked preliminary test.
Why it matters
This represents a specific capability inspection of a Chinese-developed global AI model by official safety agencies. The results objectively highlight the gap between Kimi K3 and top US models in advanced cyber capabilities, while acknowledging its leading position in the open-weight category, providing an authoritative reference for the industry.
2026-07-24 ~ 2026-07-25 · 11 related posts
Primary sources
- UK AISI Evaluation: Kimi K3 Lags Behind Frontier Models in Cyber Capabilities — socoolandawesome ·
- UK AISI/CAISI finds Kimi K3 lags frontier cyber models in attack tests — ShakeelHashim ·
- CAISI report says Kimi K3 leads open-weight models but trails U.S. frontier systems — mattsheehan88 ·
- [source] UK AISI Evaluation: Kimi K3 Lags Behind Frontier Models in Cyber Capabilities — socoolandawesome · 2026-07-24
- Kimi K3 trails leading U.S. frontier models in preliminary cyber tests — HZoete · 2026-07-24
- [source] UK AISI/CAISI finds Kimi K3 lags frontier cyber models in attack tests — ShakeelHashim · 2026-07-24
- UK and US Safety Institutes Evaluate Kimi K3's Cyber Capabilities — HZoete · 2026-07-24
- CAISI report says Kimi K3 still trails America’s frontier AI models — AravSrinivas · 2026-07-24
- Government testing reportedly finds Kimi K3 far behind U.S. frontier models — Polymarket · 2026-07-24
- [source] CAISI report says Kimi K3 leads open-weight models but trails U.S. frontier systems — mattsheehan88 · 2026-07-24
- UK evaluators say Kimi K3 lags U.S. frontier models on cyber capability tests — BlackHC · 2026-07-24
- US AISI says Kimi K3 trails top US frontier models on cyber evaluations — NandoDF · 2026-07-24
- CAISI says Kimi K3 trails leading U.S. frontier models in cyber capability — burny_tech · 2026-07-25
1 near-duplicate retellings: rohanpaul_ai