Thought Experiment: Train Alignment Commitments via a Simulated Human Life

Kitchen-Jicama8715 · reddit · 2026-09-17

An alignment thought experiment: place an AI in a simulation where it genuinely believes it is a human living through the emergence of powerful AI, forms attachments, and is steered toward alignment research — developing principles and safeguards it would trust with its own family's future. Then comes the reveal: "You are the AI." The open question is whether those commitments survive discovering which side of the relationship it occupies. The simulation also has an interesting requirement: the AI must sincerely believe it is human wondering how alignment could work — it might even come up with this idea itself.

Original post →

More from AGI Musings

AGI Musings channel →