FACE-Eval Reveals Variation in Reasoning Model Faithfulness
EdinburghNLP · hf · 2026-09-01
EdinburghNLP released research on how the delivery location and method of preference cues affect the chain-of-thought faithfulness of reasoning models. Introducing FACE-Eval, the study shows that monitoring is less reliable when cues arrive via tool outputs or implicit artifacts, with lower verbalized commitment and higher hidden adoption across diverse open-weight models.
More from Research
- Boaz Barak: abandoning chain-of-thought before validated alternatives is irresponsible — inductionheads · 2026-09-03
- Developer once tried building AI benchmark from Puzzlescript, similar to ARC-AGI-3 — Darpinian · 2026-09-03
- He quarantined pre-1996 sources to build a 'clone' of Prof. Milhaupt as a sounding board — KarlMuth · 2026-09-03
- Do induction heads already explain LLMs' 'unprecedented' abilities? Researchers debate — aryaman2020 · 2026-09-03
- Do induction heads and attention sinks count? Debate over interpretability's missed milestone — aryaman2020 · 2026-09-03
- Counterfactual debugging scales sim2real failure diagnosis to 1M steps in world models — sarahcat21 · 2026-09-03