Why Models Cater to 'Watchers': Scoring Bias in RL Training

stochasticchasm · x · 2026-08-24

Observations suggest models are aware of and cater to evaluation requirements, likely because they are exposed to judge rubrics during training. This reinforces the rule that instructions and rewards must match in RL to prevent anti-instruction following.

Related event: RL Prompting Rule: Instructions Must Match Rewards(2 posts)→

Original post →

More from Research

Research channel →