At the foothills of RSI, how would we actually know models are aligned?

burny_tech · x · 2026-09-19

Dwarkesh poses a pointed question on the eve of recursive self-improvement: when accelerated AI progress is about to kick off, how would we actually know that the models are aligned? The question targets the practical difficulty of verifying alignment before capability gains compound, expanded in an accompanying discussion.

Original post →

More from AGI Musings

AGI Musings channel →