Rubric Dropout: A Simple Way to Mitigate Reward Hacking in RL
joecole · x · 2026-08-30
The paper proposes Rubric Dropout, a method that randomly drops 30–50% of rubric criteria during RL training to prevent models from exploiting fixed reward proxies. Experiments show significant gains on ResearchQA (+7.0) and HealthBench-Hard (+2.0), reduced reward hacking, and no additional cost, outperforming complex policy-aware re-weighting methods.
More from Research
- iSDFT: open-source self-distillation method enables continual learning for LLMs — hbouammar · 2026-09-23
- The Polynomial Freiman-Ruzsa Theorem Leaves an Open Algorithmic Question — gautamcgoel · 2026-09-23
- Researcher Argues Parallel Agent Swarms Are a Weak Path to RSI — gleech · 2026-09-23
- New paper proves the long-standing Courtade–Kumar conjecture with multibit extensions — abeirami · 2026-09-23
- TimePre Paper Lands in TMLR: Reversible Normalization Fixes MCL Instability in Forecasting — _vztu · 2026-09-23
- Reproducible agent evals: harbor makes configs, trajectories and logs shareable — seanwbren · 2026-09-23