Alignment Theory Lacks Research on Rationalization

geoffreyirving · x · 2026-07-08

The author believes alignment theory is still missing a crucial puzzle piece: how to best understand "rationalization"—the process of guessing an answer first and providing justifications post-hoc.

They express hope that multiple teams will investigate this issue from different directions using various tools.

Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →