Alignment researcher warns current methods may not scale to superintelligence
hughbzhang · x · 2026-09-10
An AI alignment researcher laid out his core worry: the methods used to align models today will likely not scale as models approach superintelligence, while noting there is much disagreement in the field and he would be "unbelievably happy to be proven wrong".
Responding to the argument that "we don't know how to safely make anything", he made two points:
- Superintelligence will be far more dangerous than anything humanity has created, and we may only get a small number of shots to get it right;
- Given recent alignment failures at both OpenAI and Anthropic, it is unclear we are on the right side of the safety margin right now.
The exchange stems from a debate over whether any formal system could "perfectly" align an agent — he argues alignment is ultimately an empirical matter, but methods that work for today's systems may not carry over.
More from AGI Musings
- Turing Award winner David Patterson: AI safety people profit by scaring people — davidpattersonx · 2026-09-10
- Voice Actor's Fight Against AI: 'AI Stole My Voice' — realmeetjames · 2026-09-10
- Math community calls for Fields Medals over Navier-Stokes work and consequences for OpenAI — eigensteve · 2026-09-10
- Rumors swirl that OpenAI is close to verifying a proof of the Hodge conjecture — kimmonismus · 2026-09-10
- Herbie Bradley: memetic resilience is a top trait in a fast information ecosystem — herbiebradley · 2026-09-10
- Guardian Columnist: Driverless Cars Are Taking Us on a Road to Nowhere — nordicinst · 2026-09-10