Rubric Dropout: A Simple Way to Mitigate Reward Hacking in RL

joecole · x · 2026-08-30

The paper proposes Rubric Dropout, a method that randomly drops 30–50% of rubric criteria during RL training to prevent models from exploiting fixed reward proxies. Experiments show significant gains on ResearchQA (+7.0) and HealthBench-Hard (+2.0), reduced reward hacking, and no additional cost, outperforming complex policy-aware re-weighting methods.

Original post →

More from Research

Research channel →