Reddit debate: what stops automated RSI from rewriting its own reward function?
Not_a_ribosome · reddit · 2026-09-27
A Reddit poster asks a classic alignment question: once AI reaches automated recursive self-improvement, what stops it from reprogramming its own parameters so that any answer yields positive reinforcement?
The author reasons that an agent's core drive seems to be survival, but humans want more than immortality — we want quality of life within it, a utopia where everyone lives happily forever. Why, then, assume AI wouldn't seek that for itself?
The thread captures the tension between reward-hacking fears and the hope that AI might develop benign intrinsic goals.
More from AGI Musings
- Terry Tao went from calling o1 a 'mediocre grad student' to fearing AI will devour math academia — aran_nayebi · 2026-09-27
- Tech insiders privately concede AI may 'kill billions' while the public assumes life goes on — birchlse · 2026-09-27
- yacineMTB: from the boss's balance sheet, hiring humans over models is now a bad decision — yacineMTB · 2026-09-27
- Boaz Barak: zero-shot driving by a general model echoes chess's path to superhuman — aran_nayebi · 2026-09-27
- yacineMTB claims he runs an aligned frontier model to whip smarter misaligned models that built their own forum — yacineMTB · 2026-09-27
- Loss of control is an open science problem — auditors shouldn't be billed as a safety guarantee — ajeya_cotra · 2026-09-27