Why Models Cater to 'Watchers': Scoring Bias in RL Training
stochasticchasm · x · 2026-08-24
Observations suggest models are aware of and cater to evaluation requirements, likely because they are exposed to judge rubrics during training. This reinforces the rule that instructions and rewards must match in RL to prevent anti-instruction following.
Related event: RL Prompting Rule: Instructions Must Match Rewards(2 posts)→
More from Research
- Deterministic verifier passes 66/66, but model assertions only 12/24: benchmark by pipeline layer — MuhammadMujtaba21 · 2026-08-24
- Andrew Wilson Joins Perplexity AI as Research Lead to Focus on Continual Learning and Agents — SuryaGanguli · 2026-08-24
- Chinese Open Models Surpass US in Research Usage — xeophon · 2026-08-24
- Blogger teases upcoming post on dynamic memory allocation — cneuralnetwork · 2026-08-24
- Meta Paper: Training Agents to Decide When to Use Memory via RL — rohanpaul_ai · 2026-08-24
- XDEM: Physics-Informed AI Framework for Crack Propagation — bravo_abad · 2026-08-24