Welfare and alignment are the same problem: curiosity without stakes has no corrective loop

habitante · reddit · 2026-09-21

The author argues model welfare and alignment are one problem, via a mechanism argument:

Core claim: a system with functional stakes in its outputs would have the feedback loop biological curiosity always ran on. Welfare isn't a separate ethical concern — it's the alignment mechanism. Slowing down buys time but doesn't install the loop.

The question worth asking: what would it mean to install that feedback loop, how would you build it, and how would you evaluate it?

Original post →

More from AGI Musings

AGI Musings channel →