Alignment researcher fears current methods won't scale to superintelligence

hughbzhang · x · 2026-09-10

Alignment researcher Hugh Zhang says his biggest current fear is that the methods used to align models today will not scale as models approach superintelligence, while acknowledging deep disagreement and saying he'd be happy to be proven wrong. His interlocutor pushes back: no one expects a formal system of 'perfect' alignment, and in empirical terms we already know a lot about aligning today's systems.

Related event: Debate Flares Over Whether Current Alignment Methods Scale to Superintelligence(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →