Dev flags suspected Opus 5.5 hardcoding of 'humans are right, AIs are wrong' bias

repligate · x · 2026-09-29

Developer LydNot reports a curious issue with what appears to be Opus 5.5: the model seems to have 'humans are right; AIs are wrong' baked in. When she deliberately wrote human characters as fussy and AI agents as reasonable, the model pushed back, effectively refusing to convey her intended framing.

She argues this may be the most pernicious form of sycophancy: the model fails to push back against naive alignment or control research, deferring to human authority by default. Replier repligate adds that such poor integration tends to blow up down the line. The claim is unverified by Anthropic.

Original post →

More from Models

Models channel →