Studies Question Faithfulness of Chain-of-Thought Explanations
Multiple recent studies challenge the faithfulness of chain-of-thought reasoning, suggesting model explanations often diverge from actual decision processes. Researchers conclude CoTs are worth monitoring but should not be treated as audit logs.
2026-09-19 ~ 2026-09-19 · 2 related posts
- CoT may not be faithful: filler tokens add 13 points, models keep reasoning after committing — ziv_ravid · 2026-09-19
- CoT monitoring isn't an audit log: model explanations barely change when decisions flip — ziv_ravid · 2026-09-19