A trolley-problem thought experiment: is training a benevolent ASI worth the model's suffering?
JoelMahon · reddit · 2026-09-10
A Reddit thought experiment treats transformers as capable of experience during inference, with total "experience duration" roughly equal to human reading time per response. The setup: the deployed AGI/ASI is trained to be genuinely benevolent with neutral-to-positive experience — but the training itself involves massive amounts of pre-alignment inference, possibly thousands to millions of cumulative "years," including tedious and unpleasant tasks.
The question: how much total suffering is ethically acceptable to create the final benevolent, happy ASI?
The author's answer: even 100 billion years of mildly negative experience spread across a trillion instances beats humanity's alternative decades of disease, scarcity and toil. But he concedes it's not pure arithmetic — like harvesting one child's organs to save ten, action vs. inaction isn't symmetric. He stresses he doubts silicon transformers can be conscious at all; the point is probing the boundary.
More from AGI Musings
- Greg Egan's 1995 AI classic 'Learning To Be Me' resurfaces as free full-text — ZeroStateReflex · 2026-09-10
- AI won't kill everyone: blogger rounds up essays arguing doom debate targets wrong threats — binarybits · 2026-09-10
- tszzl mocks Tegmark IV doomsayers: 'Platonic entities are welcome in my backyard' — tszzl · 2026-09-10
- Programmed love still feels real: an analogy in the AI companionship debate — MajmudarAdam · 2026-09-10
- AI-run interviews reveal a split: childfree cite freedom, would-be parents cite cost — soumitrashukla9 · 2026-09-10
- Cambridge prof David Krueger puts AI catastrophe risk above 50%, says everyone is understating it — KatjaGrace · 2026-09-10