CoT may not be faithful: filler tokens add 13 points, models keep reasoning after committing

ziv_ravid · x · 2026-09-19

A thread surveying recent papers that cast doubt on whether chain-of-thought is a faithful record of the computation producing the answer.

Implication: be skeptical of CoT as a safety-monitoring channel.

Related event: Studies Question Faithfulness of Chain-of-Thought Explanations(2 posts)→

Original post →

More from Safety

Safety channel →