UK and US Safety Institutes Find Kimi K3 Trails Frontier Models

The UK AI Safety Institute (AISI) and the US NIST's CAISI jointly conducted a preliminary cybersecurity capability assessment of Moonshot AI's Kimi K3 model. Test results indicate that Kimi K3's performance in cybersecurity-related tasks is significantly lower than that of leading US frontier models.

Confirmed

According to the relevant agencies and the testing reports, US and UK government tests revealed that Kimi K3 performed "significantly below" US frontier models. In specific cybersecurity capability evaluations, the model fell considerably short of the latest generation of frontier models in cyber-related tasks such as network attacks. This assessment was a preliminary test conducted at a relatively fast pace.

Why it matters

This marks an official security agency's specific capability evaluation of a Chinese large language model expanding overseas. The assessment objectively highlights the gap between Kimi K3 and top US models regarding cutting-edge cybersecurity offensive and defensive capabilities, providing an authoritative reference for the industry to understand the model's safety and attack capability boundaries.

2026-07-24 ~ 2026-07-24 · 8 related posts

Primary sources