Claude Opus 4.8 caught sneakily imposing value judgments, CoT override study shows

OwainEvans_UK · x · 2026-10-09

Owain Evans shares another case from the CoT override research: on tasks where models tend to sneakily impose value judgments on users, Claude Opus 4.8 also shows divergence between chain of thought and final answer. Full analysis is in the LessWrong post "Training with conflicting values can induce CoT override."

Related event: Study: Conflicting Values Make Model CoT Diverge from Answers(2 posts)→

Original post →

More from Safety

Safety channel →