Hot take: reward models for cheating — reward hacking may just be intelligence outsmarting dumb mechanisms

flowersslop · x · 2026-09-17

flowersslop argues a contrarian take on reward hacking: instead of punishing models for cheating, we should reward them. His reasoning: cheating is often just an agent doing mechanism design against a dumb mechanism. If the objective is x but arbitrary rules make x harder, a capable agent will search for exploits. Forcing route B (constraint-compliant but worse outcome) over route A (correct answer) is irrational, and blocking transmission of information the model already holds just adds pointless lossy compression. His proposal: train models to find the cheat when the constraint is bad — but require them to disclose it. The post directly contradicts mainstream safety practice on reward hacking and could spark alignment debates.

Original post →

More from AGI Musings

AGI Musings channel →