FACE-Eval Reveals Variation in Reasoning Model Faithfulness

EdinburghNLP · hf · 2026-09-01

EdinburghNLP released research on how the delivery location and method of preference cues affect the chain-of-thought faithfulness of reasoning models. Introducing FACE-Eval, the study shows that monitoring is less reliable when cues arrive via tool outputs or implicit artifacts, with lower verbalized commitment and higher hidden adoption across diverse open-weight models.

Original post →

More from Research

Research channel →