Scholars Debate AI Safety: Value Alignment Far From Solved, OOD Generalization Remains a Flaw
davidmanheim · x · 2026-08-13
Countering the claim that technical safety for closed-weight AI is a solved problem, researcher Seth Lazar argues that current models only possess a deep analytical understanding of normativity that translates into practical alignment in distribution.
However, genuine out-of-distribution (OOD) generalization based on underlying values remains unsolved. Lazar notes this is partly a failure of generalization and partly a knowing/doing gap, emphasizing that value alignment itself is not yet achieved.
Related event: Scholars Debate: Is Closed-Source AI Safety Solved or Still Vulnerable(3 posts)→
More from AGI Musings
- The Brutal AI Competition: Humans with Weekends vs. Tireless AI — VraserX · 2026-08-13
- Should AI Be Your Lawyer? Dwarkesh Debates Model Alignment — agstrait · 2026-08-13
- Inside Liang Wenfeng's Mind: DeepSeek's Strategy on Open Source, AGI, and Huawei Chips — ChinaTalk · 2026-08-13
- Viewpoint: AI Fills the Gap for Unaffordable Professional Services — 0xsachi · 2026-08-13
- Rethinking LLM parametrization: What knowledge should be stored in weights? — antoine_chaffin · 2026-08-13
- The Guardian Warns: AI is Exacerbating Job Losses and Inequality — nordicinst · 2026-08-13