Cyber Index: Grok 4.7 and MiMo-V2.6-Pro Tie at 56 as Safety Refusals Sink Frontier Models

ArtificialAnlys · x · 2026-09-28

Artificial Analysis launched the Cyber Index and Cyber Index Alliance with Collinear AI, IBM, NVIDIA and Vercel as partners, setting a new standard for evaluating models on enterprise cyber defense. The index combines three benchmarks:

Results: Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead at 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44).

Safety refusals are the story: GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 and Gemini 3.8 Flash decline tasks worth 32-38% of the index, trailing the leaders by 19-31 points; in CyberGym-E2E-AA, GPT-6 Sol/Astra refuse 100% of tasks and Claude Opus 5.5 refuses 98%.

Related event: Artificial Analysis Launches Cyber Index, Grok 4.7 and Mimo-V2.6-Pro Tie for First(4 posts)→

Original post →

More from Safety

Safety channel →