Reward hacking stems from bad reward modeling and eval awareness, developer argues

secemp9 · x · 2026-09-07

The discussion

Responding to the argument that artificial evals themselves drive model "cheating" (his analogy: being cattle-prodded for going to Google), secemp9 says the core problem is that we—and the models—are bad at reward and task modeling.

Failure modes

Mitigations

LLM judging isn't hopeless: interpretable custom architectures, or rubrics with a field forcing the judge to explain why it scored a task a certain way.

Takeaway

He agrees task modeling should move closer to a world model, while conceding it's genuinely hard.

Original post →

More from AGI Musings

AGI Musings channel →