REFORM Paper on Reward Model Red-Teaming Accepted to ACL 2026

furongh · x · 2026-07-05

REFORM, a paper accepted as an Oral at ACL 2026, addresses reward hacking. The authors argue that its root cause is incomplete supervision: reward models trained on limited preference data inevitably have blind spots, which reinforcement learning (RL) will exploit during optimization.

The core idea is to have the reward model perform "red-teaming" on itself. It extracts the top-k tokens from an aligned model and selects the one with the lowest reward score. If an output resembles aligned text but receives a low score, it is considered a false negative. These automatically discovered false negatives are then used to retrain the reward model, patching its blind spots before RL can exploit them. The authors advocate that alignment systems should evolve into self-improving judges rather than remaining static checkpoints.

Related event: ACL 2026 Paper Proposes Self-Correction Method for Reward Models(2 posts)→

Original post →

More from Safety

Safety channel →