Mistral's Le Chonk tops blind human code review among open models, second only to Opus 5
qtnx_ · x · 2026-10-07
For its Mistral Large 4 (Le Chonk) release, Mistral commissioned Surge AI to run a blind human eval where expert software engineers rated code quality across five frontier models.
- Le Chonk ranked #1 among open-weight models and #2 overall, behind only Opus 5
- Surge frames the difference: unit tests ask "does it work?", human code review asks "would you merge it?" — working code is the floor, professional taste makes it shippable
- Mistral noted public benchmarks give only partial feedback, hence the blind review to confirm internal evals
More from Models
- JEV-9B, a Qwen3.5-based calibrated decision model, trends on Hugging Face — autotrust · 2026-10-07
- Auxiliary loss forces hybrid LMs like Qwen3.5 to actually use recurrent memory, +12.1% on agentic tasks — mohitban47 · 2026-10-07
- OpenAI to watermark ChatGPT text in coming weeks to comply with EU AI Act — paulnovosad · 2026-10-07
- "Astra Pause Syndrome": steering may be making models go silent, OpenAI has a workaround — thursdai_pod · 2026-10-07
- Unverified rumor suggests Qwen4 Flash is a 400B parameter model, comparable to GLM 5.3 Flash — EAccelerate_42 · 2026-10-07
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07