Stanford paper's Counterfactual Simulation Training boosts CoT faithfulness monitoring by 35 points
a_karvonen · x · 2026-09-04
Peter Hase and Christopher Potts introduce Counterfactual Simulation Training (CST) in a new arXiv paper: CoTs are rewarded when they let a simulator accurately predict the model's outputs on counterfactual inputs, improving Chain-of-Thought faithfulness.
Two settings:
- Cue-based counterfactual CoT monitoring, detecting spurious-feature reliance, reward hacking, and sycophancy;
- Generic model-based counterfactual simulation encouraging more faithful, generalizable reasoning.
Key results (models up to 235B):
- CST improves monitor accuracy on cue-based counterfactuals by 35 accuracy points and simulatability on generic counterfactuals by 2 points;
- It outperforms prompting baselines;
- Rewriting unfaithful CoTs with an LLM is 5x more efficient than RL alone;
- Faithfulness gains don't generalize to dissuading cues;
- Larger models aren't more faithful out of the box, but benefit more from CST.
More from Safety
- NYT Reveals the Hugging Face Hack Involved 700 AIs 'Sacrificing' Each Other — dylfreed · 2026-09-04
- 'Develop or deploy' wording means using today's AI could carry 20-year prison risk — kevinnbass · 2026-09-04
- Who actually runs adversarial testing in the Agent Development Lifecycle? — Specialist-Bee9801 · 2026-09-04
- LAUSD quietly imposes districtwide moratorium on student generative AI use on school devices — chrismattmann · 2026-09-04
- 1Password Responds to a User's Disappointment, and the Full Exchange Is Documented — backlit4034 · 2026-09-04
- Merkl launches a cryptographic notary layer for AI agents, with open-source SDK — Efficiency_Positive · 2026-09-04