UK/US Agencies Evaluate Kimi K3: Lags Behind Frontier Models in Cyber Capabilities

A joint preliminary assessment by UK AISI and US CAISI evaluated the cyber capabilities of Moonshot's Kimi K3. The findings reveal that Kimi K3 significantly lags behind US frontier models in cyber-related tasks, though it leads among open-weight models.

Confirmed

According to information shared by the agencies and related sources, the tests found that Kimi K3's performance is "significantly lower" than US frontier models. In specific cyber capability evaluations, it fell short of the latest generation of frontier models in offensive cyber tasks. For instance, in the ExploitBench test, Kimi K3 stopped at an average of 17 steps; by comparison, the frontier model Mythos Preview achieved Arbitrary Code Execution (ACE) in 18 out of 41 tasks. Additionally, charts from the Commerce / CAISI-NIST report indicate that the comparison included other models like GLM-5.2. The evaluation was noted as a fast-tracked preliminary test.

Why it matters

This represents a specific capability inspection of a Chinese-developed global AI model by official safety agencies. The results objectively highlight the gap between Kimi K3 and top US models in advanced cyber capabilities, while acknowledging its leading position in the open-weight category, providing an authoritative reference for the industry.

2026-07-24 ~ 2026-07-25 · 11 related posts

Primary sources

1 near-duplicate retellings: rohanpaul_ai