Claude Opus 5 tops a cybersecurity benchmark but becomes noisier when it overworks

Thom_Wolf · x · 2026-07-29

A cybersecurity benchmark run on Claude Opus 5 found that it uncovers more vulnerabilities than other frontier models, slightly ahead of GPT-5.6 Sol.

The authors also say Opus 5 is noticeably more “hyperactive” than earlier generations: that extra effort helps it find more bugs, but it makes the results noisy and less stable.

Original post →

More from Models

Models channel →