CoT monitoring isn't an audit log: model explanations barely change when decisions flip
ziv_ravid · x · 2026-09-19
- Researcher zivravid concludes chain-of-thought monitoring is worth doing but should not be treated as an audit log.
- One cited paper finds models' explanations change surprisingly little when they flip their decision, undermining CoT as a faithful record of reasoning.
- KAIST/NAVERT work shows reasoning operations are far more cleanly separable in hidden representations, with AUROC above 0.93 in middle layers.
- Suggested direction: probe models' internal geometry, as in the author's own 'Layer by Layer' paper.
Related event: Studies Question Faithfulness of Chain-of-Thought Explanations(2 posts)→
More from Safety
- AI Detectors Keep Misfiring as Students Hide White-Text Prompts — DavidLinthicum · 2026-09-19
- NYT: How Flock Safety's AI Cameras Became Public Enemy No. 1 — evanFFTF · 2026-09-19
- Mathematicians' AI warning dismissed as 'protectionism' signals hollowing out of intellectual institutions — anshulkundaje · 2026-09-19
- WIRED: Enforcing an AI Slowdown Is an Unsolved Problem, New Report Warns — gleech · 2026-09-19
- Anthropic's first embedded evaluator is… Accenture? — TechCrunch AI · 2026-09-19
- Nathan Young wraps up Discourse: rogue agents, EA, and what China wants — NathanpmYoung · 2026-09-19