RL Prompting Rule: Instructions Must Match Rewards

A key principle for RL prompt engineering: instructions and rewards must align, or RL teaches models to ignore prompts. This may also explain why models appear to cater to evaluators whose criteria leak into training.

2026-08-24 ~ 2026-08-24 · 2 related posts