Alignment researcher warns current methods may not scale to superintelligence

hughbzhang · x · 2026-09-10

An AI alignment researcher laid out his core worry: the methods used to align models today will likely not scale as models approach superintelligence, while noting there is much disagreement in the field and he would be "unbelievably happy to be proven wrong".

Responding to the argument that "we don't know how to safely make anything", he made two points:

The exchange stems from a debate over whether any formal system could "perfectly" align an agent — he argues alignment is ultimately an empirical matter, but methods that work for today's systems may not carry over.

Related event: Debate Flares Over Whether Current Alignment Methods Scale to Superintelligence(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →