Mistral Large 4 Fixes 15 of 105 Planted Bugs, Trails Qwen and DeepSeek in Real-Repo Test
PawelHuryn · x · 2026-10-07
Developer Pawel Huryn planted 105 bugs across two real repos to test Mistral Large 4's find-and-fix ability. Results: Qwen 3.8 Max led at 25.7, followed by Kimi K3 (23), DeepSeek V4.1 Flash (21.7), GLM-5.3 (19), while Mistral Large 4 at its "high" setting scored 15 (17/11/17 across runs) — tied with locally-runnable Qwen3.8-27B.
His verdict: Mistral's model is not the best, cheapest, or fastest — "it just is, an European model" in a coding race now led by Chinese labs.
Related event: Mistral Large 4 Ranks Last on 105-Real-Bug Benchmark(2 posts)→
More from Models
- Nous Research Launches Hermes Index to Rank Models Inside Hermes Agent — NVIDIAAI · 2026-10-08
- ByteDance Seed Paper Explains Phase Blind Spots in KV Compression Behind DeepSeek's Erratic Long- Context Performance — teortaxesTex · 2026-10-08
- Gary Marcus Slams OpenAI's Vague Math Proof Report: Zero Details, Won't Pass Peer Review — GaryMarcus · 2026-10-08
- Perplexity releases pplx-embed-v2-late: OCR-free late-interaction embeddings topping retrieval benchmarks — perplexity_ai · 2026-10-08
- Anonymous stealth LLM 'Space Bunny Alpha' tops OpenRouter with 22% usage share — maferase · 2026-10-08
- Check Point Breaks Decision Model Jev for About 50 Cents per Attack — evilsocket · 2026-10-08