Meta's Two-Stage RP-OPD + RL Beats SFT + RL on Rubric Tasks and Curbs Reward Hacking
meta · hf · 2026-10-07
Meta researchers propose a two-stage framework for open-ended tasks that can't be verified by exact outcomes: first, rubric-privileged on-policy distillation (RP-OPD) uses rubrics as teacher-only context for dense token-level supervision, then rubric-based RL pushes past the distillation plateau. On HealthBench, ResearchQA, and RubricHub Science with open-weight models, RP-OPD + RL scores highest among compared post-training methods and shows little reward hacking, while the SFT + RL baseline increasingly earns high rewards by merely claiming rubric compliance. The takeaway: use rubrics to guide on-policy distillation before rubric-based RL.
More from Research
- Perplexity releases open multimodal late-interaction embedding family with SOTA results — antoine_chaffin · 2026-10-08
- User challenges frontier AI labs to drop 700+ cancer-cure papers as the true AGI milestone — BLUECOW009 · 2026-10-08
- Paper: Vibe Coding Kills Open Source as AI-Recommended Repos Lose Stars — soumitrashukla9 · 2026-10-08
- Scaling AI-guided experiments beats scaling biological data, researcher argues at ICML — anshulkundaje · 2026-10-08
- MA-BC: Provably Efficient Multi-Objective Imitation Learning from Heterogeneous Experts — Yossarian_1234 · 2026-10-08
- Hide Model Names From Agents: Labels Cost 55% More Tokens and Drop Success to 81% — alex_verem · 2026-10-08