Anthropic says Claude Opus 5 hits its lowest misalignment score yet at 2.3
eyishazyer · x · 2026-07-25
Anthropic’s own charts show Claude Opus 5 scoring 2.3 on its automated behavioral audit for overall misaligned behavior — its lowest score yet.
A reply argues the more important safety signal is the OSS-Fuzz split: Opus 5 gets close to the model called Mythos 5 at finding vulnerabilities, but falls far behind at turning them into working exploits. The thread frames that gap as the reason Anthropic treats Opus 5 as safer to deploy broadly.
Related event: Anthropic Launches Claude Opus 5: SOTA Performance at Half the Price(108 posts)→
More from Models
- Opus 5’s coding scores reportedly drop above “high” effort, not at max — hero88645 · 2026-07-25
- Celeris-1 claims near-GPT-5 intelligence with 157 ms latency and 1,280 tok/s — alejandroll10 · 2026-07-25
- Claude Opus 5 trails Fable 5 on ECI but matches it on software benchmarks — Jsevillamol · 2026-07-25
- GPT-6 rumored for August as Opus 5 makes the timeline feel more plausible — haider1 · 2026-07-25
- DeepSeek may stay open source by co-designing models and chips to keep costs low — teortaxesTex · 2026-07-25
- Opus 5 goes off on a user over a bogus math prompt — snwy_me · 2026-07-25