Ilya's four-year-old 'teach the AGI to love' resurfaces as a value alignment proposal
morqon · x · 2026-09-10
A thread resurfaces Ilya Sutskever's four-year-old tweet "Gotta teach the AGI to love," arguing it remains increasingly relevant to value alignment.
The author argues nature already solved alignment: a mother stays aligned with her offspring via (1) modeling the child's own value function, (2) emotional motivation that rewards her when the child's motivation function is rewarded, and (3) a real-time feedback loop from the child's responses that penalizes model inaccuracies. These three network properties, he says, are love — and alignment research should copy them.
More from AGI Musings
- Notion CEO-shared take: own your context, rent the intelligence — ivanhzhao · 2026-09-10
- Gary Marcus makes the case for boycotting generative AI over reliability — GaryMarcus · 2026-09-10
- Anthropic researcher rebuts the claim that averted AI risk means worriers were fools — geoffreyirving · 2026-09-10
- tszzl: Aligning near-human AI is a reasonable proxy for studying superintelligence alignment — tszzl · 2026-09-10
- Anthropic researcher puts >10% odds on AI killing all humans; skeptics quip 'we can unplug it' — inductionheads · 2026-09-10
- AI researcher on academia: when intelligence is universal, curiosity is what endures — kwangmoo_yi · 2026-09-10