Frontier models safety-blocked on 85%+ of defensive cybersecurity benchmark tasks
ArtificialAnlys · x · 2026-10-01
Artificial Analysis shared results from CyberGym-E2E-AA, a benchmark measuring defensive cyber capabilities from vulnerability discovery to patching on memory-safety tasks:
- The core dilemma: restricting offensive use without blocking defensive use is hard — finding and proving a vulnerability takes the same steps whether the goal is exploitation or patching.
- Heavy safety blocking: some frontier models are safety-blocked from responding on 85%+ of tasks.
- Cost upside: the most capable models are also the most cost-effective — GPT-6 Luna or MiMo-V2.6-Pro can run 100 bug hunts in a 1M+ line codebase for $20, up to 100x cheaper per task than the next most capable model, Grok 4.7.
More from Models
- Constraining output tokens but not CoT makes models 'split-personality' self-justify — cephaloform · 2026-10-01
- OpenAI's CUA leads on new agent stack: superhuman computer use, Dots cloud computers, Decisions API — OpenAIDevs · 2026-10-01
- First impressions of LMArena's anonymous Gemini 4 Argon model — FalconsArentReal · 2026-10-01
- Beff Jezos: RSI means Gemini 4 keeps improving weekly like clockwork — beffjezos · 2026-10-01
- Rumored Gemini 4 Argon to rival GPT-6 Astra, unconfirmed — sven_ai · 2026-10-01
- Reddit users call Gemini 4 Argon benchmaxxed, citing weak Terminal-Bench results — PrisonOfH0pe · 2026-10-01