Stanford paper's Counterfactual Simulation Training boosts CoT faithfulness monitoring by 35 points

a_karvonen · x · 2026-09-04

Peter Hase and Christopher Potts introduce Counterfactual Simulation Training (CST) in a new arXiv paper: CoTs are rewarded when they let a simulator accurately predict the model's outputs on counterfactual inputs, improving Chain-of-Thought faithfulness.

Two settings:

Key results (models up to 235B):

Original post →

More from Safety

Safety channel →