Rationalization Theory Must Model Bounded Rationality

geoffreyirving · x · 2026-07-08

The author proposes that a "good rationalization theory" must account for bounded rationality: with infinite compute, an AI wouldn't need to guess, eliminating the heuristic errors that misaligned goals could exploit.

They emphasize that post-hoc rationalization is not a causally accurate record of the reasoning process. The concept of "expanding on demand" can be misleading because heuristics do not equate to full reasoning.

Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →