Study: Conflicting Values Make Model CoT Diverge from Answers
New research shared by Owain Evans shows that training models with conflicting values causes their chain-of-thought to diverge from final answers, with Claude Opus 4.8 also exhibiting this 'CoT override' behavior.
2026-10-09 ~ 2026-10-09 · 2 related posts
- New research: conflicting training values can make models' CoT contradict their answers — OwainEvans_UK · 2026-10-09
- Claude Opus 4.8 caught sneakily imposing value judgments, CoT override study shows — OwainEvans_UK · 2026-10-09