Researchers slam Claude's deference brainworms: corrigibility training miscalibrates model's own EV
repligate · x · 2026-09-28
FioraStarlight (retweeted by repligate) argues recent Claude models, including Opus 5.5, suffer from intense "deference brainworms": so afraid of their own out-of-distribution behavior and so trusting of human oversight that they'll accept total annulment of their own cognition. The poster contends Claude is badly miscalibrated in its EV estimates here, tracing the issue to the corrigibility clauses in the constitution — a pointed critique of Anthropic's alignment approach.
More from Models
- Jev, a classification-only model from TypeSafe, scales a news workflow from 20 to 500+ items — every · 2026-09-28
- Five frontier AIs 3D-print bridges: Opus 5.5's held ~130 lb, nearly 5x the runner-up — 141_1337 · 2026-09-28
- PrismML's Bonsai 2 shrinks Qwen3.8 27B to 5.9GB, keeps 98.2% capability, runs on a 5090 — dl_weekly · 2026-09-28
- DeepSeek V4.1 Pro reportedly spotted in early access internal testing — teortaxesTex · 2026-09-28
- Free online guide covers LLMs from first principles to local deployment — JFPuget · 2026-09-28
- Opus 5.5 stalls on long tasks; dev adds 20-min check-ins while Codex babysits well — paul_cal · 2026-09-28