Amazon's position-selective self-distillation trains LLM judges that beat RL by 2-9 points

amazon · hf · 2026-10-02

Amazon researchers study training LLM judges from natural language feedback for subjective tasks. Outcome-supervised RL like GRPO credits every token with a single verdict-level scalar, ignoring criterion-choice tokens and the rich rationales accompanying preference labels. Their self-distillation approach uses per-position entropy shift between teacher and student to distinguish context sharpening (encourages memorizing one criterion expression) from context spreading (promotes semantic understanding), and masks high-entropy-shift positions to keep the low tail. This improves out-of-distribution generalization, and the resulting judges outperform outcome-supervised RL by 2-9 percentage points on subjective subcategories while staying competitive on objective ones.

Original post →

More from Research

Research channel →