Alignment debate: SGD's inductive biases may not generalize to human values

akbirthko · x · 2026-09-25

In an alignment debate, Kaushik Reddy argues against the view that capabilities and alignment generalize together: inductive biases that select well for capabilities (e.g. SGD selecting for simplest programs) fail to generalize to human values. A concrete technical counterpoint to the "stronger models are safer" assumption.

Original post →

More from AGI Musings

AGI Musings channel →