Kimi K3 trails Mythos on cyber-range tasks in a benchmark chart shared online
petrusenko_max · x · 2026-07-25
A benchmark chart comparing Mythos Preview, Kimi K3, and GLM-5.2 claims Kimi K3 is still behind Mythos on cyber-range tasks that resemble real-world offensive utility.
The chart breaks performance into categories such as full exploit, general primitives, V8 primitives, bug reproduction, and coverage. Kimi K3 scores 0 on the first two categories, trails Mythos on V8 primitives and bug reproduction, and only matches the others on coverage, leading the author to argue that China is still at least six months behind the U.S. in this space.
Related event: Kimi K3 Cybersecurity Eval Sparks Debate: Scores 32.2%(5 posts)→
More from Models
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11