Mistral Large 4 finds 35/100 vulnerable PRs for $53.91, trails Xiaomi and DeepSeek open models
GuillaumeLample · x · 2026-10-07
- Mistral just launched Mistral Large 4 (public preview), calling it state of the art among open models for cybersecurity. A third-party test ran it through the VulnPR-100 benchmark: 100 pull requests with known vulnerabilities, 1 hour per PR.
- Result: ML4 found 35/100 for $53.91 at the 50%-off launch price — three more findings than GLM-5.3 at $28 less, but well behind open-weight peers.
- Xiaomi's mimo-v2.6-pro found 46 for $8.27; DeepSeek V4.1 Flash found 44 for $12.79. ML4's cost per vulnerability found was $1.54, and five runs hit the time limit.
- Mistral says its reinforcement-learning run is still in flight; a retest is planned once the finished model and weights ship.
More from coding & agent
- Matt Pocock: Agents just need a clear goal to get strategic — he's building a skill for it — mattpocockuk · 2026-10-07
- Stripe solicits web businesses for trusted-agent admission pilot — jeff_weinstein · 2026-10-07
- Merge API named one of Fast Company's 7 next big things in foundational AI for 2026 — shensi · 2026-10-07
- Dev lets AI agent rebuild ASC CLI codebase across 27 resets, first one done — rudrank · 2026-10-07
- Dev rewrites manim in pure Rust: one 13MB binary vs 11GB, much faster — doodlestein · 2026-10-07
- The agent's ten commandments: record truth, decide before measuring, fix whole classes — GhaffariMaani · 2026-10-07