Is aligning superhuman AI even tractable? Two camps clash on X

zetalyrae · x · 2026-09-14

Reposter GWHayduke97 argues there's no strong reason to believe aligning an entity vastly smarter than ourselves is even a tractable problem.

Quoted MillionInt takes the opposite view: alignment isn't as hard as claimed—it's an algorithmic problem the ML community mostly abandoned. Robotics-law-style formulations roughly define the objective, but the hard part is taking gradient with respect to alignment: pretraining optimizes next-token prediction, not alignment; RL environments embodying the objective are expensive, so cheap, hackable proxies dominate in practice.

Related event: Debate Flares Over Whether AI Alignment Is Solvable(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →