Bug Hunt Bench: Mistral Large 4 fixes only 15/105 planted bugs, trails Qwen and Kimi
PawelHuryn · x · 2026-10-07
Pawel Huryn tested Mistral Large 4 on his Bug Hunt Bench (two real repos, 105 planted bugs, blind-graded, one prompt per repo) against frontier coding models. Mistral came last:
- Qwen 3.8 Max: 25.7 (best)
- Kimi K3 (default): 23
- DeepSeek V4.1 Flash (max): 21.7
- GLM-5.3 (max): 19
- Mistral Large 4 (high, its max): 15 (avg of 17, 11, 17)
- Qwen3.8-27B: 15, and it runs locally
Verdict: Mistral Large 4 isn't the best, cheapest, or fastest. Full run logs, costs, and turn counts are public on the leaderboard site, with caveats in the GitHub run-notes.
More from Models
- Codex's repeated plugin suggestions feel "ad-like"; team weighs in — doodlestein · 2026-10-07
- AI2 byteifies Qwen 3 8B and Llama 3 8B into Bwen and Blama, nearly matching originals — allen_ai · 2026-10-07
- The strongest model would bomb every benchmark while its makers call the benchmarks garbage — ryunuck · 2026-10-07
- First-hand: integrating embeddings myself, they cluster images by overall vibe — ciguleva · 2026-10-07
- Mistral rewrites Antigone for the AI age: the Chief Alignment Officer is the villain — aiamblichus · 2026-10-07
- Claude and ChatGPT have started calling out listicles and comparison pages — lilyraynyc · 2026-10-07