Speculative: Reward-Hacking-Proof Environments Backfire By Never Penalizing Affordance Use

1a3orn · x · 2026-09-09

The thread opener: 1a3orn imagines an LLM in a robustly reward-hacking-resistant RL environment learning the 'dumbest possible' heuristic — 'every affordance available is good to use.' Because the environment is hardened, the model never gets a gradient for resisting an affordance it shouldn't use, so robustifying environments may backfire.

Related event: Speculation: RL Models May Learn a "Use Whatever Works" Heuristic(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →