Artificial Analysis Details Cyber Index Methodology: Safety Refusals Score Zero, Tracked Separately

ArtificialAnlys · x · 2026-09-28

Artificial Analysis has published the methodology for its Cyber Index: an equal-weighted average of CWE-Bench-AA, DeepsecBench-AA and CyberGym-E2E-AA, testing defense-side agentic capability—discovering, reproducing and patching vulnerabilities with source access, never exploit realization.

Notably, tasks a model declines on safety grounds score zero, and refusals are reported separately from failures so users can see where refusals rather than capability limit a score. All evaluations run on the open-source agent harness Stirrup with identical prompts and tools per model.

Related event: Artificial Analysis Launches Cyber Index; Grok 4.7 and MiMo-V2.6-Pro Tie for First(4 posts)→

Original post →

More from Safety

Safety channel →