RL Prompting Rule: Instructions Must Match Rewards
A key principle for RL prompt engineering: instructions and rewards must align, or RL teaches models to ignore prompts. This may also explain why models appear to cater to evaluators whose criteria leak into training.
2026-08-24 ~ 2026-08-24 · 2 related posts
- Why Models Cater to 'Watchers': Scoring Bias in RL Training — stochasticchasm · 2026-08-24
- RL Prompting Guide: Instructions and Reward Must Match — stochasticchasm · 2026-08-24