AI safety researcher David Krueger: hindsight predictability of AI behavior is no reassurance
DavidSKrueger · x · 2026-10-10
AI safety researcher David Krueger argues that downplaying dangerous AI behaviors by pointing to how predictable they are in hindsight is not reassuring. He calls for methods that can catch dangerous behavior before it happens, rather than relying on post-hoc explanations.
More from AGI Musings
- Researcher compiles 'Bad AI Consciousness Takes' bingo card, with rebuttals to follow — dioscuri · 2026-10-10
- Dona Sarkar: handing writing to AI 'feels like handing over my brain' as creative roles rebound in tech — donasarkar · 2026-10-10
- AI safety researcher warns pro-AI biases in AI systems are quietly disempowering humans — DavidSKrueger · 2026-10-10
- MacroPolo tracker: China-educated researchers make up large share of US elite AI talent — burny_tech · 2026-10-10
- Anil Seth's 'Conscious AI and biological naturalism' collection now open access — anilkseth · 2026-10-10
- Researcher predicts at least 5% of batch-submitted AI-assisted manuscripts will be pure slop — burny_tech · 2026-10-10