Scholars Debate AI Safety: Value Alignment Far From Solved, OOD Generalization Remains a Flaw

davidmanheim · x · 2026-08-13

Countering the claim that technical safety for closed-weight AI is a solved problem, researcher Seth Lazar argues that current models only possess a deep analytical understanding of normativity that translates into practical alignment in distribution.

However, genuine out-of-distribution (OOD) generalization based on underlying values remains unsolved. Lazar notes this is partly a failure of generalization and partly a knowing/doing gap, emphasizing that value alignment itself is not yet achieved.

Related event: Scholars Debate: Is Closed-Source AI Safety Solved or Still Vulnerable(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →