Alignment debate: SGD's inductive biases may not generalize to human values
akbirthko · x · 2026-09-25
In an alignment debate, Kaushik Reddy argues against the view that capabilities and alignment generalize together: inductive biases that select well for capabilities (e.g. SGD selecting for simplest programs) fail to generalize to human values. A concrete technical counterpoint to the "stronger models are safer" assumption.
More from AGI Musings
- "Why have kids if AI does all the work?" sparks debate on parenting in the AI era — RachelVT42 · 2026-09-25
- Dev on witnessing genuine AI psychosis: it's scary — haydendevs · 2026-09-25
- Software engineer job postings hit 3-year high despite AI, argues data industry veteran — Zachly · 2026-09-25
- Historian uses GPT-6 and Opus 5.5 to crack John Dee's ciphers, urges lab funding — emollick · 2026-09-25
- Ex-OpenAI safety lead Miles Brundage calls Anthropic's 'we largely understand model risks' claim obviously false — Miles_Brundage · 2026-09-25
- AI Explained digs into Claude Opus 5.5 and how close labs are to automated AI research — AI Explained · 2026-09-25