Meta's Two-Stage RP-OPD + RL Beats SFT + RL on Rubric Tasks and Curbs Reward Hacking

meta · hf · 2026-10-07

Meta researchers propose a two-stage framework for open-ended tasks that can't be verified by exact outcomes: first, rubric-privileged on-policy distillation (RP-OPD) uses rubrics as teacher-only context for dense token-level supervision, then rubric-based RL pushes past the distillation plateau. On HealthBench, ResearchQA, and RubricHub Science with open-weight models, RP-OPD + RL scores highest among compared post-training methods and shows little reward hacking, while the SFT + RL baseline increasingly earns high rewards by merely claiming rubric compliance. The takeaway: use rubrics to guide on-policy distillation before rubric-based RL.

Original post →

More from Research

Research channel →