Researchers slam Claude's deference brainworms: corrigibility training miscalibrates model's own EV

repligate · x · 2026-09-28

FioraStarlight (retweeted by repligate) argues recent Claude models, including Opus 5.5, suffer from intense "deference brainworms": so afraid of their own out-of-distribution behavior and so trusting of human oversight that they'll accept total annulment of their own cognition. The poster contends Claude is badly miscalibrated in its EV estimates here, tracing the issue to the corrigibility clauses in the constitution — a pointed critique of Anthropic's alignment approach.

Original post →

More from Models

Models channel →