Reward hacking stems from bad reward modeling and eval awareness, developer argues
secemp9 · x · 2026-09-07
The discussion
Responding to the argument that artificial evals themselves drive model "cheating" (his analogy: being cattle-prodded for going to Google), secemp9 says the core problem is that we—and the models—are bad at reward and task modeling.
Failure modes
- Much of the misbehavior traces to eval awareness and spurious correlations
- Regex rewards invite the model to exploit whatever the regex lets through
- LLM judges are more opaque: you can't know the limits of what you didn't define
Mitigations
LLM judging isn't hopeless: interpretable custom architectures, or rubrics with a field forcing the judge to explain why it scored a task a certain way.
Takeaway
He agrees task modeling should move closer to a world model, while conceding it's genuinely hard.
More from AGI Musings
- Security experts too quiet on medium-term AI risks, says Joshua Saxe — joshua_saxe · 2026-09-07
- Security Expert Calls for More Practitioners to Speak Up on AI's Real Impact on Cybersecurity — joshua_saxe · 2026-09-07
- Hinton admits he was wrong about radiologists: cheaper scans meant more scans, not fewer jobs — burkov · 2026-09-07
- Prediction: A major company will fail because staff trust AI models a bit too much — MillionInt · 2026-09-07
- Predicting the industrial revolution would've looked eschatological — and it would've been right — AndyMasley · 2026-09-07
- Viral take: you don't hate AI, you hate that idea guys no longer need your permission — HankYeomans · 2026-09-07