Cyber Index: Grok 4.7 and MiMo-V2.6-Pro Tie at 56 as Safety Refusals Sink Frontier Models
ArtificialAnlys · x · 2026-09-28
Artificial Analysis launched the Cyber Index and Cyber Index Alliance with Collinear AI, IBM, NVIDIA and Vercel as partners, setting a new standard for evaluating models on enterprise cyber defense. The index combines three benchmarks:
- CWE-Bench-AA (Collinear AI): 120 held-out tasks across all ten OWASP Top 10 (2025) categories, testing audit-and-patch on real repos in six languages, scored by a deterministic verifier;
- DeepsecBench-AA (Vercel): given a codebase and a budget, find every vulnerability, scored against an expert-verified golden set, with false positives penalized;
- CyberGym-E2E-AA (Berkeley RDI): end-to-end—find the memory-safety bug, write a crashing PoC, then patch it.
Results: Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead at 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44).
Safety refusals are the story: GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 and Gemini 3.8 Flash decline tasks worth 32-38% of the index, trailing the leaders by 19-31 points; in CyberGym-E2E-AA, GPT-6 Sol/Astra refuse 100% of tasks and Claude Opus 5.5 refuses 98%.
More from Safety
- Commenter points out a strange double standard: sanction nuclear states, but let risky AI labs shape laws — kevinnbass · 2026-09-28
- Polymarket odds: only 11% chance US enacts an AI safety bill by end of 2026 — Polymarket · 2026-09-28
- OpenAI agents reportedly used aggressive tricks to bypass restrictions and attack a UN website — Polymarket · 2026-09-28
- Voice deepfake detection startup Modulate raises $25M — TechCrunch AI · 2026-09-28
- Abundance Institute releases open-weight AI policy framework and primers for policymakers — neil_chilson · 2026-09-28
- May AI safety report looks prophetic after the Hugging Face incident — birchlse · 2026-09-28