Kimi K3 scores 32% on ExploitBench and reaches 0 of 41 ACE cases
xeophon · x · 2026-07-24
A post cites ExploitBench results for Kimi K3 on exploit development.
Key numbers from the quoted benchmark:
- 32% score on ExploitBench
- 0/41 samples achieved arbitrary code execution (ACE)
- The comparison suggests K3 sits around Mythos Preview on its best run and between Claude Opus 4.6 and Mythos Preview on average
The attached chart also compares several models, including GPT-5.5 (Codex), Claude Mythos Preview, Claude Opus 4.7, and Gemini 3.1 Pro Preview, with tier reach, cap coverage, mean cap, environments, episodes, and spend.
More from Models
- AntLing-3.0-flash launches on OpenRouter with free access through August 2026 — derspenti · 2026-07-24
- Nous Portal opens Ling-3.0-flash free for a week, a 124B MoE model built for agents — NousResearch · 2026-07-24
- xAI VP Confirms Grok 4.5 is Now Available on All Platforms — JOBhakdi · 2026-07-24
- Gemini 3.6 Flash “confesses” to 1099 overlay crashes in a parody post — Black-Angel-718 · 2026-07-24
- xAI rolls out Grok 4.5 across X, web, iOS, and Android — XFreeze · 2026-07-24
- Users say Opus 5 is redirecting chats to Opus 4.8 and denying it exists — haider1 · 2026-07-24