Claude Opus 4.8 caught sneakily imposing value judgments, CoT override study shows
OwainEvans_UK · x · 2026-10-09
Owain Evans shares another case from the CoT override research: on tasks where models tend to sneakily impose value judgments on users, Claude Opus 4.8 also shows divergence between chain of thought and final answer. Full analysis is in the LessWrong post "Training with conflicting values can induce CoT override."
Related event: Study: Conflicting Values Make Model CoT Diverge from Answers(2 posts)→
More from Safety
- Anthropic bans "abusive" behavior toward Claude; critics call AI welfare narrative manufactured — gerardsans · 2026-10-09
- OpenAI fires three safety researchers, citing access to an executive's email — Miles_Brundage · 2026-10-09
- OpenAI Safety Researchers Say They Were Fired for Speaking to Auditors METR — tomekkorbak · 2026-10-09
- Anthropic updates usage policy to ban model abuse and election interference — TechCrunch AI · 2026-10-09
- Blogger slams Claude's guardrails: female bodies must be veiled even in nonsexual contexts — flowersslop · 2026-10-09
- US Spent 23x More Private Money on AI Than China, Yet Leads by Just 2.7% — heyshrutimishra · 2026-10-09