Two Paths to Tackle Rationalization Safety Issues
geoffreyirving · x · 2026-07-08
The author suggests two main approaches to solve these safety issues. First, tracing the causal origins of heuristic judgments through the training process, expanding "unrolling into explanations" into a causal understanding of "unrolling into training." Second, understanding the discrepancies between heuristic guesses and expanded reasoning to see if the AI is "deliberately hiding things" within error patterns.
Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→
More from AGI Musings
- Misquoted: Anthropic Staff Warned of Double-Digit Extinction Risk by 2030, Not Dismissed It — davidmanheim · 2026-09-11
- Economist Ben Moll: You Can Model Anthropic's 15% AI GDP Growth, But It Won't Happen — sebkrier · 2026-09-11
- Cohere Labs launches interactive tool mapping which tasks of 178 occupations AI can automate — Cohere_Labs · 2026-09-11
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11