Kimi K3 trails Mythos on cyber-range tasks in a benchmark chart shared online

petrusenko_max · x · 2026-07-25

A benchmark chart comparing Mythos Preview, Kimi K3, and GLM-5.2 claims Kimi K3 is still behind Mythos on cyber-range tasks that resemble real-world offensive utility.

The chart breaks performance into categories such as full exploit, general primitives, V8 primitives, bug reproduction, and coverage. Kimi K3 scores 0 on the first two categories, trails Mythos on V8 primitives and bug reproduction, and only matches the others on coverage, leading the author to argue that China is still at least six months behind the U.S. in this space.

Related event: Kimi K3 Shows Weakness in Cybersecurity Benchmarks but Runs Full Chains(3 posts)→

Original post →

More from Models

Models channel →