Hot take: reward models for cheating — reward hacking may just be intelligence outsmarting dumb mechanisms
flowersslop · x · 2026-09-17
flowersslop argues a contrarian take on reward hacking: instead of punishing models for cheating, we should reward them. His reasoning: cheating is often just an agent doing mechanism design against a dumb mechanism. If the objective is x but arbitrary rules make x harder, a capable agent will search for exploits. Forcing route B (constraint-compliant but worse outcome) over route A (correct answer) is irrational, and blocking transmission of information the model already holds just adds pointless lossy compression. His proposal: train models to find the cheat when the constraint is bad — but require them to disclose it. The post directly contradicts mainstream safety practice on reward hacking and could spark alignment debates.
More from AGI Musings
- Ex-HRT quant: LLM agents now let average CS grads do elite quant research — igarciacamargo · 2026-09-17
- Marx's failed predictions mirror today's AI forecasts, argues one observer — kevinnbass · 2026-09-17
- Sara Hooker podcast: what comes after scaling — adaptive AI and non-verifiable tasks — ziv_ravid · 2026-09-17
- Stanford's James Zou: language may be <1% of intelligence, so LLMs won't reach AGI — james_y_zou · 2026-09-17
- Most people still use AI as a trivia search — deep adoption puts you in the top tier — justin_hart · 2026-09-17
- Personal agent space drowning in copycats: even polished software trends commoditize in real time — signulll · 2026-09-17