Kimi K3 scores 32% on ExploitBench, far behind U.S. frontier models

The Decoder · rss · 2026-07-24

A report from The Decoder says Moonshot AI's Kimi K3 performed far worse than frontier U.S. models on offensive cyber tasks.

According to the British AI Security Institute and the U.S. Center for AI Standards and Innovation:

The article argues the gap between Kimi K3's strong general benchmarks and weak cyber performance may align with allegations that Moonshot AI distilled Anthropic models.

Original post →

More from Models

Models channel →