Two Paths to Tackle Rationalization Safety Issues

geoffreyirving · x · 2026-07-08

The author suggests two main approaches to solve these safety issues. First, tracing the causal origins of heuristic judgments through the training process, expanding "unrolling into explanations" into a causal understanding of "unrolling into training." Second, understanding the discrepancies between heuristic guesses and expanded reasoning to see if the AI is "deliberately hiding things" within error patterns.

Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →