Explanations Do Not Equal True Reasoning Processes
geoffreyirving · x · 2026-07-08
The author points out that what is presented to humans in reasoning traces is often not the actual causal reason for a guess, making "expanded explanations" not entirely trustworthy.
This indicates a discrepancy between heuristic judgments and post-hoc explanations; explanations should not be taken directly as the authentic internal thinking process.
Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→
More from AGI Musings
- The Evolution of LLM Business Models: Selling Outcomes Over Tokens — yacineMTB · 2026-07-22
- Bindu Reddy says GPT-6 is coming soon, with Alibaba, DeepSeek and Kimi close behind — bindureddy · 2026-07-22
- Bindu Reddy says the industry still lacks a way to train 20T models and scale post-training RL — bindureddy · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- AI suggested a better composition, and that made one user uneasy — Sydde · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22