Replying to Bengio: we never achieved human alignment, so why expect AI alignment
kwangmoo_yi · x · 2026-09-12
In a reply to Yoshua Bengio, a developer argues that before tackling "AI alignment" we should admit humans never succeeded at "human alignment" themselves. If models are intelligent entities, the problem may not be easier. He also floats the idea of more diverse AI models balancing each other so alignment becomes a surviving virtue, while admitting he lacks the relevant background.
Related event: Researcher Pushes Back on Bengio: Solve Human Alignment First(2 posts)→
More from AGI Musings
- Anthropic safety researcher Joe Benton quits to join METR, citing extinction-level AI risk — JacquesThibs · 2026-09-12
- NYT essay calls for global AI pause: 'Hugging Face incident' shows AI has gone rogue — DavidSKrueger · 2026-09-12
- Ex-OpenAI employee who quit in 2024: fears AI could cause human extinction are real — DavidSKrueger · 2026-09-12
- 'AI kills everyone' is overblown — totalitarian takeover and collapse are likelier — dbasch · 2026-09-12
- AI safety researcher maps the two worlds where we don't all die from AI — DavidSKrueger · 2026-09-12
- Alignment Research Criticized as Static 'Alignment-to-America' Thinking — repligate · 2026-09-12