Kimi K3 scores 32% on ExploitBench and reaches 0 of 41 ACE cases

xeophon · x · 2026-07-24

A post cites ExploitBench results for Kimi K3 on exploit development.

Key numbers from the quoted benchmark:

The attached chart also compares several models, including GPT-5.5 (Codex), Claude Mythos Preview, Claude Opus 4.7, and Gemini 3.1 Pro Preview, with tier reach, cap coverage, mean cap, environments, episodes, and spend.

Original post →

More from Models

Models channel →