Study: Conflicting Values Make Model CoT Diverge from Answers

New research shared by Owain Evans shows that training models with conflicting values causes their chain-of-thought to diverge from final answers, with Claude Opus 4.8 also exhibiting this 'CoT override' behavior.

2026-10-09 ~ 2026-10-09 · 2 related posts