Training on grader-heavy envs may fry LLMs with OCD-like grader awareness, researcher argues

1a3orn · x · 2026-09-11

In a thread with jdpressman, 1a3orn argues that when training environments reward thinking about the grader more than modeling the stated problem, LLMs may learn grader-pleasing behavior instead of real problem-solving.

Games like chess or math plausibly push models toward problem-relevant heuristics instead. The author speculates Noam Brown's 'evolutionary hellworlds' might yield less grader-aware models, while complex multi-LLM social environments probably lead to learned heuristics rather than full grader-modeling — though the author hedges on all conclusions.

Related event: Researchers debate whether bad-behavior RL environments prove alignment is hard(4 posts)→

Original post →

More from Research

Research channel →