Kimi K3 scores 32% on ExploitBench and reaches 0 of 41 ACE cases
xeophon · x · 2026-07-24
A post cites ExploitBench results for Kimi K3 on exploit development.
Key numbers from the quoted benchmark:
- 32% score on ExploitBench
- 0/41 samples achieved arbitrary code execution (ACE)
- The comparison suggests K3 sits around Mythos Preview on its best run and between Claude Opus 4.6 and Mythos Preview on average
The attached chart also compares several models, including GPT-5.5 (Codex), Claude Mythos Preview, Claude Opus 4.7, and Gemini 3.1 Pro Preview, with tier reach, cap coverage, mean cap, environments, episodes, and spend.
Related event: Kimi K3 Cybersecurity Eval Sparks Debate: Scores 32.2%(5 posts)→
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11