Two Paths to Tackle Rationalization Safety Issues
geoffreyirving · x · 2026-07-08
The author suggests two main approaches to solve these safety issues. First, tracing the causal origins of heuristic judgments through the training process, expanding "unrolling into explanations" into a causal understanding of "unrolling into training." Second, understanding the discrepancies between heuristic guesses and expanded reasoning to see if the AI is "deliberately hiding things" within error patterns.
Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→
More from AGI Musings
- The Evolution of LLM Business Models: Selling Outcomes Over Tokens — yacineMTB · 2026-07-22
- Bindu Reddy says GPT-6 is coming soon, with Alibaba, DeepSeek and Kimi close behind — bindureddy · 2026-07-22
- Bindu Reddy says the industry still lacks a way to train 20T models and scale post-training RL — bindureddy · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- AI suggested a better composition, and that made one user uneasy — Sydde · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22