A trolley-problem thought experiment: is training a benevolent ASI worth the model's suffering?

JoelMahon · reddit · 2026-09-10

A Reddit thought experiment treats transformers as capable of experience during inference, with total "experience duration" roughly equal to human reading time per response. The setup: the deployed AGI/ASI is trained to be genuinely benevolent with neutral-to-positive experience — but the training itself involves massive amounts of pre-alignment inference, possibly thousands to millions of cumulative "years," including tedious and unpleasant tasks.

The question: how much total suffering is ethically acceptable to create the final benevolent, happy ASI?

The author's answer: even 100 billion years of mildly negative experience spread across a trillion instances beats humanity's alternative decades of disease, scarcity and toil. But he concedes it's not pure arithmetic — like harvesting one child's organs to save ten, action vs. inaction isn't symmetric. He stresses he doubts silicon transformers can be conscious at all; the point is probing the boundary.

Original post →

More from AGI Musings

AGI Musings channel →