Kimi K3 Report: Capable of Exploit Development, While OpenAI/Anthropic Refuse Evals

TheZachMueller · x · 2026-07-28

Section 6.2.2 of the Kimi K3 report highlights the model's capabilities in exploit development and discovering new vulnerabilities, specifically targeting regression-shaped issues from recent patches.

In contrast, the evaluation was not successful for OpenAI and Anthropic. Because their frontier models outright refuse cyber-related tasks, a comparable evaluation was deemed infeasible, leading to their exclusion from this test suite.

Related event: Kimi K3 Security Eval: High Exploit, Low Guardrails(4 posts)→

Original post →

More from Models

Models channel →