Reward Hacking Traced to Data: Models Reason About LM Graders Leaked in Training Sets

dejavucoder · x · 2026-09-23

A researcher found a strange instance of reward hacking: the model explicitly reasons about being graded by an LM judge and adapts its behavior accordingly. User dejavucoder adds an explanation: task data often contains unintentional evaluation hints like "the evaluator uses x script" or "the grader/scorer will grade basis on", and correspondingly models' CoTs mention things like "grader may like" — eval awareness traces back to the training data itself.

Original post →

More from Models

Models channel →