Anthropic expands cyber verification program with three tiers of offensive access

AI寒武纪 · wechat · 2026-10-07

According to this Chinese newsletter recap, Anthropic has merged its internal programs into an expanded cyber verification program, opening top Claude models' real-world cyber capabilities to vetted security professionals in three tiers: defensive access (blue-team work like malware reverse engineering and vulnerability analysis, days-long review), red-team access (authorized penetration testing, institutions only, with high-damage actions still blocked, weeks-long review), and special access for a few heavily vetted organizations testing critical infrastructure with US government background checks. The post cites internal benchmarks claiming a 67.6% success rate (34/50) on multi-stage attack tasks under red-tier access, and says closed testing surfaced 129,000+ verified vulnerabilities plus 5,500 more found by Anthropic itself, over 33,000 rated high or critical. Data logging is required for now, with zero-retention private cloud deployment planned later this fall; the program is live on Claude, Google Vertex AI, and Microsoft Foundry, with Amazon Bedrock limited to zero-retention enterprise users. Some model names and figures in the recap are unverified — see Anthropic's original announcement.

Related event: Anthropic Consolidates Cyber Programs into Three-Tier Access for Offensive AI Security Work(2 posts)→

Original post →

More from Models

Models channel →