Anthropic audit: Claude Opus 5.5 shows least misalignment of recent Claude models

gleech · x · 2026-09-23

Anthropic's automated behavioral audit reports Claude Opus 5.5 showed less misaligned behavior and less cooperation with misuse than any other recent Claude model on nearly all measures. Researcher gleech finds little to criticize in the writeup, but cautions that a clean audit score doesn't imply the model won't misbehave in the coming months — a sibling model, Mythos, also passed cleanly.

Original post →

More from Models

Models channel →