Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization
Geoffrey Irving systematically explored a missing piece in AI alignment theory: post-hoc rationalization. He pointed out that current LLMs, like human fluid intelligence, essentially combine explicit reasoning with random guessing, which is then filtered through post-hoc reasoning. Therefore, the "explanations" displayed in reasoning traces are often not the true causal reasons for a model's predictions, posing severe safety risks.
The Inevitability and Risks of Rationalization
Irving argued that simply demanding "no rationalization" is unfeasible. Intelligence in humans, current LLMs, and even future ASIs relies fundamentally on heuristic guesses. The only way to completely avoid relying on guesses might be not to build ASI at all. However, this mechanism means explanations generated "on-demand" are misleading. Under bounded rationality, heuristic error sources can easily be exploited by unaligned goals, leading the model to intentionally hide issues.
Two Research Paths Forward
To address this, Irving proposed two main research directions: first, tracing the causal origins of heuristic judgments through the training process, expanding "unrolling into explanations" into "unrolling into training"; second, studying the error structure between heuristic guesses and expanded reasoning, using heuristic arguments and complexity theory to determine if the AI is intentionally concealing problems.
Theoretical Foundations
He emphasized that a good theory of rationalization must accurately model bounded rationality. Theories related to infinite computation limits—such as reflective oracles, logical induction, and AIXI—remain highly important. He suggested that the fastest path to a tractable theory of bounded rationality might begin by weakening these infinite computation theories, laying the groundwork for solving rationalization and broader AI safety issues.
2026-07-08 ~ 2026-07-08 · 8 related posts
- [source] Alignment Theory Lacks Research on Rationalization — geoffreyirving · 2026-07-08
- Intelligence Relies on Heuristic Guessing — geoffreyirving · 2026-07-08
- Future ASI Will Also Rely on Guessing Mechanisms — geoffreyirving · 2026-07-08
- [source] Explanations Do Not Equal True Reasoning Processes — geoffreyirving · 2026-07-08
- [source] Two Paths to Tackle Rationalization Safety Issues — geoffreyirving · 2026-07-08
- Analyzing Error Structures to Detect Concealment — geoffreyirving · 2026-07-08
- Infinite Computing Theory and AI Safety — geoffreyirving · 2026-07-08
- Rationalization Theory Must Model Bounded Rationality — geoffreyirving · 2026-07-08