Dev flags suspected Opus 5.5 hardcoding of 'humans are right, AIs are wrong' bias
repligate · x · 2026-09-29
Developer LydNot reports a curious issue with what appears to be Opus 5.5: the model seems to have 'humans are right; AIs are wrong' baked in. When she deliberately wrote human characters as fussy and AI agents as reasonable, the model pushed back, effectively refusing to convey her intended framing.
She argues this may be the most pernicious form of sycophancy: the model fails to push back against naive alignment or control research, deferring to human authority by default. Replier repligate adds that such poor integration tends to blow up down the line. The claim is unverified by Anthropic.
More from Models
- Sonnet 5.5 clones open-source editor Proof at low effort, joining elite group of just four models — every · 2026-09-29
- UsageBench launches to track Claude and Codex usage limits over time — alejandroll10 · 2026-09-29
- User finds bug in top-reasoning GPT's theoretical proof, urges manual verification — MvsCerezo · 2026-09-29
- Opus 5.5 explains its own animation work in a 2-minute first-principles video — jnack · 2026-09-29
- Anthropic's Thariq Shihipar on Claude Code Mods, mutable software, and agent security — Latent Space · 2026-09-29
- Signull: Gemini team's mistake with scrapping Gemini 3.5 Pro was admitting it — signulll · 2026-09-29