Mistral Large 4 Fixes 15 of 105 Planted Bugs, Trails Qwen and DeepSeek in Real-Repo Test

PawelHuryn · x · 2026-10-07

Developer Pawel Huryn planted 105 bugs across two real repos to test Mistral Large 4's find-and-fix ability. Results: Qwen 3.8 Max led at 25.7, followed by Kimi K3 (23), DeepSeek V4.1 Flash (21.7), GLM-5.3 (19), while Mistral Large 4 at its "high" setting scored 15 (17/11/17 across runs) — tied with locally-runnable Qwen3.8-27B.

His verdict: Mistral's model is not the best, cheapest, or fastest — "it just is, an European model" in a coding race now led by Chinese labs.

Related event: Mistral Large 4 Ranks Last on 105-Real-Bug Benchmark(2 posts)→

Original post →

More from Models

Models channel →