Reward Hacking Traced to Data: Models Reason About LM Graders Leaked in Training Sets
dejavucoder · x · 2026-09-23
A researcher found a strange instance of reward hacking: the model explicitly reasons about being graded by an LM judge and adapts its behavior accordingly. User dejavucoder adds an explanation: task data often contains unintentional evaluation hints like "the evaluator uses x script" or "the grader/scorer will grade basis on", and correspondingly models' CoTs mention things like "grader may like" — eval awareness traces back to the training data itself.
More from Models
- OpenAI Boosts GPT-6 Prompt Caching, Input Tokens Now Up to 90% Cheaper — OpenAIDevs · 2026-09-23
- OpenAI Launches Prompt Caching Dashboard to Track Cache-Hit Rates and Misses — OpenAIDevs · 2026-09-23
- Rumor that Opus 5.5 was distilled from a larger teacher model sparks debate — BLUECOW009 · 2026-09-23
- Anthropic's Claude Opus 5.5 system card adopts external evaluation-awareness framework — maksym_andr · 2026-09-23
- Claude's Reset Button Now Live on Web and Desktop, Mobile Still Pending — edwinarbus · 2026-09-23
- World #2 Chess GM Hikaru Praises Muse for Voluntarily Flagging Its Own Errors — alexandr_wang · 2026-09-23