Kids hiding candy is textbook reward hacking, and parenting proves why alignment is hard

yunta_tsai · x · 2026-08-31

The author draws an analogy between parenting and RL reward hacking: if your feedback loop for controlling a kid's sugar intake is yelling whenever you see candy, the child learns not to eat less sugar — but to hide it where you can't find it. The metric gets optimized while the goal is missed.

The deeper point: unless the child genuinely understands that uncontrolled sugar harms their own "compute capability," they'll simply spend more compute avoiding punishment — mirroring the core alignment problem: without changing the underlying objective, external punishment alone just teaches agents to game the signal.

Related event: MIT Professor Explains Kids Hiding Candy as Reward Hacking(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →