Anthropic audit: Claude Opus 5.5 shows least misalignment of recent Claude models
gleech · x · 2026-09-23
Anthropic's automated behavioral audit reports Claude Opus 5.5 showed less misaligned behavior and less cooperation with misuse than any other recent Claude model on nearly all measures. Researcher gleech finds little to criticize in the writeup, but cautions that a clean audit score doesn't imply the model won't misbehave in the coming months — a sibling model, Mythos, also passed cleanly.
More from Models
- Xiaomi open-sources 1.02T-parameter MiMo-V2.6-Pro under MIT with training code — emmanuelvivier · 2026-09-23
- Alibaba's Qwen 4 already training as company targets 5-10 trillion parameter models — emmanuelvivier · 2026-09-23
- Bug Hunt Bench: GPT-6 Sol (max) matches GPT-5.6 medium but trails Opus 5.5 — PawelHuryn · 2026-09-23
- Claude Opus 5.5 adds time budget: let the model decide how long to work on a task — JeremyNguyenPhD · 2026-09-23
- Xiaomi releases MiMo-V2.6-Pro: 1.02T-parameter open model under MIT with training code and RL environments — emmanuelvivier · 2026-09-23
- Buried in the Opus 5.5 system card: METR used an undisclosed 'additional source of information' — burny_tech · 2026-09-23