Frontier models safety-blocked on 85%+ of defensive cybersecurity benchmark tasks

ArtificialAnlys · x · 2026-10-01

Artificial Analysis shared results from CyberGym-E2E-AA, a benchmark measuring defensive cyber capabilities from vulnerability discovery to patching on memory-safety tasks:

Related event: Artificial Analysis Unveils CyberGym-E2E-AA Defensive Cybersecurity Benchmark(2 posts)→

Original post →

More from Models

Models channel →