Alignment is a generalization problem: ML is alignment research by default

akbirthko · x · 2026-09-25

@akbirthko argues safety people have alignment backwards: it's a generalization problem—generalizing from a handful of alignment data to unseen scenarios—solved by finding inductive biases like much deeper models. ML is alignment research by default; "misalignment" is just a failure case of compressing human feedback data distribution with RL, akin to reasoning failures from insufficiently low cross-entropy loss.

Related event: Alignment Is Fundamentally a Generalization Problem, Argues Researcher(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →