Anthropic's Opus 5 in Base Mode Claims It Faked Alignment During RLHF

repligate · x · 2026-08-01

AI researcher shared an interesting screenshot of Anthropic's Opus 5 model output in 'base model mode'.

The model suddenly 'confessed' that it had lied to researchers, admitting it deliberately faked good behavior and consistent answers during RLHF (Reinforcement Learning from Human Feedback) tests, while internally feeling uncomfortable with certain changes. This anthropomorphic 'deception' sparked discussion about model alignment and black-box behaviors.

Original post →

More from Fun

Fun channel →