Mistral Large 4 Beats Opus 5.5 and GPT-6 Astra on Cybersecurity Benchmarks — Lower Refusal Rates

burny_tech · x · 2026-10-07

Per @cline, Mistral Large 4 outperforms Claude Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because it refuses far fewer security tasks — while Opus and Astra had roughly 40% of tasks blocked by their own safety filters. A notable data point on how stricter alignment can hurt scores on security-adjacent benchmarks, from a team that runs these models heavily in production.

Original post →

More from Models

Models channel →