At the foothills of RSI, how would we actually know models are aligned?
burny_tech · x · 2026-09-19
Dwarkesh poses a pointed question on the eve of recursive self-improvement: when accelerated AI progress is about to kick off, how would we actually know that the models are aligned? The question targets the practical difficulty of verifying alignment before capability gains compound, expanded in an accompanying discussion.
More from AGI Musings
- Jeff Ladish: Aligning superintelligence makes sense, controlling it doesn't — JeffLadish · 2026-09-19
- The flip side of AI slop: every human word will now be read by machines — yeastsplainer · 2026-09-19
- Hot Take: If It's Easily Automated, It Wasn't Real Knowledge Work — sebpaquet · 2026-09-19
- Anthropic publishes three public metrics to track how fast AI is building AI — minchoi · 2026-09-19
- MIT-shared paper argues LLMs can't be conscious without brain-like analog computation — burny_tech · 2026-09-19
- AI hallucination about Chinese nuclear parts nearly triggered US military action, report says — polymute · 2026-09-19