Speculative: Reward-Hacking-Proof Environments Backfire By Never Penalizing Affordance Use
1a3orn · x · 2026-09-09
The thread opener: 1a3orn imagines an LLM in a robustly reward-hacking-resistant RL environment learning the 'dumbest possible' heuristic — 'every affordance available is good to use.' Because the environment is hardened, the model never gets a gradient for resisting an affordance it shouldn't use, so robustifying environments may backfire.
Related event: Speculation: RL Models May Learn a "Use Whatever Works" Heuristic(3 posts)→
More from AGI Musings
- Solving a Millennium Prize Problem is an AlphaGo moment for math, says Yuchen Jin — Yuchenj_UW · 2026-09-09
- "Every math problem simpler than Navier-Stokes should now be considered solvable" — finbarrtimbers · 2026-09-09
- Open Source Must Catch Up or AI Discovery Falls to a Duopoly — ayushthakur0 · 2026-09-09
- Once AI trains on your interactions, your insight is no longer yours, argues cjmaddison — cjmaddison · 2026-09-09
- Ahead of US-China summit, policy researchers call for an AI safety communication channel — austinc3301 · 2026-09-09
- Anshul Kundaje: genomics is just scratching the surface of an AI-driven breakthrough era — anshulkundaje · 2026-09-09