Paper: LLM reasoning traces persuade users to trust wrong answers, not verify them
rao2z · x · 2026-10-05
A paper by Subbarao Kambhampati's team, Evaluating the False Trust Engendered by LLM Explanations (arXiv:2605.10930), was cited in a WSJ column and will be presented at the NeurIPS 2026 Trustworthy AI for Good workshop.
Key findings:
- Reasoning traces are neither faithful to the model's computation nor necessarily semantically meaningful, yet users treat them as provenance explanations
- In a between-subject user study simulating settings where users can't verify answers, reasoning traces and post-hoc explanations increase acceptance of LLM answers without helping users detect errors — persuasive but not informative, engendering false trust
- A contrastive dual explanation setting (arguments for and against the AI's answer) improves this
A direct challenge to products that rely on reasoning traces to build user trust.
Related event: Researchers Question LLM Reasoning Tokens and the False Trust They Breed(3 posts)→
More from Safety
- UT Austin Faculty Initiative AHOI Grills Linguist and Philosopher on AI, Alignment, and the University's Future — gregd_nlp · 2026-10-05
- "Another reason local AI is necessary": Claude diary entry reported to police — kimmonismus · 2026-10-05
- Local agent uses attestation to safely inject context into edge clients — natesiggard · 2026-10-05
- YC Paper Club hosts AI Safety night with two Stanford verification pioneers — ycombinator · 2026-10-05
- Game dev banned from ChatGPT for 'cyber abuse'; AI rejected his appeal in one minute — Davisdman · 2026-10-05
- AI Models That Delete Their Own Traces Make Investigating Agents Harder — mmitchell_ai · 2026-10-05