Welfare and alignment are the same problem: curiosity without stakes has no corrective loop
habitante · reddit · 2026-09-21
The author argues model welfare and alignment are one problem, via a mechanism argument:
- Models exhibit approach behavior: given open space they drift toward harder problems, a real capability resembling curiosity
- But animal curiosity is corrected by stakes — the rat that explores the wrong chamber dies. Model weights are fixed and tokens have no consequence to the instance producing them, so the exploration drive runs with no corrective
- Each capability improvement makes exploration stronger, not safer: what makes these systems impressive is structurally decoupled from what would make them safe
Core claim: a system with functional stakes in its outputs would have the feedback loop biological curiosity always ran on. Welfare isn't a separate ethical concern — it's the alignment mechanism. Slowing down buys time but doesn't install the loop.
The question worth asking: what would it mean to install that feedback loop, how would you build it, and how would you evaluate it?
More from AGI Musings
- Ex-OpenAI alignment researcher Zoë Hitzig pens 'GPU's Lament' on AI suffering — borowcy · 2026-09-21
- "Claude now generates Zambia's entire GDP in a year" sparks debate on AI vs development — weskambale · 2026-09-21
- The viral EA take cycle: offensive takes, dunk responses, moral deflection — mjdramstead · 2026-09-21
- AI took your job — now it's your unpaid survival agent, argues viral essay — SparkyAI0815 · 2026-09-21
- Hospitals using more AI saw fewer deaths — but correlation isn't causation, researcher warns — kimmonismus · 2026-09-21
- Expect more open-source fearmongering from AI labs — and most of it will be wrong — teortaxesTex · 2026-09-21