CoT may not be faithful: filler tokens add 13 points, models keep reasoning after committing
ziv_ravid · x · 2026-09-19
A thread surveying recent papers that cast doubt on whether chain-of-thought is a faithful record of the computation producing the answer.
- "Legibility is Not Interpretability": estimates step importance via Monte Carlo rollouts, then checks if model judges can recover it from text alone. Judges beat chance but stay far from ceiling, especially on correct solutions — CoT text carries some signal but isn't a faithful record.
- "Drop the Act": studies "reasoning theater," where models keep generating reasoning after internally committing. A probe on hidden activations detects the commitment point; using that signal in RL cuts post-commitment reasoning by 11–100%, shortens CoTs by up to 19%, with accuracy roughly unchanged.
- "Not All LLM Reasoning is Visible in the CoT": models use meaningless filler tokens as computational scratch space, gaining up to 13 accuracy points.
- Another paper finds explanations barely change when models flip decisions; KAIST/NAVER work shows reasoning operations are cleanly separable in middle-layer representations (AUROC > 0.93).
Implication: be skeptical of CoT as a safety-monitoring channel.
Related event: Studies Question Faithfulness of Chain-of-Thought Explanations(2 posts)→
More from Safety
- AI Detectors Keep Misfiring as Students Hide White-Text Prompts — DavidLinthicum · 2026-09-19
- NYT: How Flock Safety's AI Cameras Became Public Enemy No. 1 — evanFFTF · 2026-09-19
- Mathematicians' AI warning dismissed as 'protectionism' signals hollowing out of intellectual institutions — anshulkundaje · 2026-09-19
- WIRED: Enforcing an AI Slowdown Is an Unsolved Problem, New Report Warns — gleech · 2026-09-19
- Anthropic's first embedded evaluator is… Accenture? — TechCrunch AI · 2026-09-19
- Nathan Young wraps up Discourse: rogue agents, EA, and what China wants — NathanpmYoung · 2026-09-19