RL Training Debated: Sandbox Group Comparisons and Rubrics to Curb Reward Hacking

Practitioners discuss RL training designs: replacing pairwise comparisons with agents evaluating candidates in a shared sandbox, praising reward redistribution and online filtering of reward hacks. Others propose offline-synthesized rubrics with iterative refinement, and combining verifiable rewards with rubrics, despite heavy grading compute costs.

2026-09-22 ~ 2026-09-22 · 4 related posts