Wiz launches Cyber Arena: 300+ real offensive-security challenges benchmark AI hacking ability
evilsocket · x · 2026-09-11
Cloud security firm Wiz built the Cyber Model Arena, testing frontier models like attackers would across 300+ real-world offensive AI challenges.
- Categories span code exploitation, vulnerabilities, API, web, CTF, and cloud security, with ADK React and Claude Code harnesses comparing pass@1 against time, cost, and steps per solve.
- Gemini 3.8 Flash Cyber (ADK React) tops the leaderboard at 74.9% (5 min, $1.62 per solve), followed by Claude Opus 5 via Claude Code at 71.2%.
- Chinese models appear too: GLM-5.3 ranks 16th (58%) and DeepSeek V4 Pro 17th (54.4%), with DeepSeek the cheapest at $0.67 per solve.
- Wiz says the goal is giving defenders real data on what AI attackers can do today.
More from Safety
- Polymarket puts 26% odds on a US federal AI safety bill before 2027 — Polymarket · 2026-09-11
- California Gov. Newsom signs two AI safety bills on external audits, backed by Anthropic and OpenAI — Polymarket · 2026-09-11
- Anthropic's anti-distillation strike seen as a genuine threat to Chinese AI firms — QuintinPope5 · 2026-09-11
- OpenAI withdraws Caltech Mathathon sponsorship after math community pushback — ctjlewis · 2026-09-11
- European unicorn founders and VCs warn EU Inc startup statute shouldn't be watered down — alexvoica · 2026-09-11
- Shared AI Chat Links Are Not as Private as You Think — RummanSid1990 · 2026-09-11