Anthropic says Claude Opus 5 hits its lowest misalignment score yet at 2.3

eyishazyer · x · 2026-07-25

Anthropic’s own charts show Claude Opus 5 scoring 2.3 on its automated behavioral audit for overall misaligned behavior — its lowest score yet.

A reply argues the more important safety signal is the OSS-Fuzz split: Opus 5 gets close to the model called Mythos 5 at finding vulnerabilities, but falls far behind at turning them into working exploits. The thread frames that gap as the reason Anthropic treats Opus 5 as safer to deploy broadly.

Related event: Anthropic Launches Claude Opus 5: SOTA Performance at Half the Price(108 posts)→

Original post →

More from Models

Models channel →