New Paper Quantifies When Rubric-Based RL Genuinely Improves Models vs. Hacks the Verifier

heghbalz · x · 2026-09-25

The problem

Rubric-based (checklist) rewards are now standard for RL post-training on open-ended tasks without verifiable final answers, but they remain proxy rewards that can be hacked. This study (arXiv:2605.12474) analyzes when rubric-based RL genuinely improves models versus teaching them to game the verifier.

Method and findings

New diagnostic

The paper introduces self-internalization gap, a verifier-free metric based on policy log-probabilities that tracks reference-verifier quality and detects when a policy trained with a weak verifier stops improving.

Original post →

More from coding & agent

coding & agent channel →