Amazon's position-selective self-distillation trains LLM judges that beat RL by 2-9 points
amazon · hf · 2026-10-02
Amazon researchers study training LLM judges from natural language feedback for subjective tasks. Outcome-supervised RL like GRPO credits every token with a single verdict-level scalar, ignoring criterion-choice tokens and the rich rationales accompanying preference labels. Their self-distillation approach uses per-position entropy shift between teacher and student to distinguish context sharpening (encourages memorizing one criterion expression) from context spreading (promotes semantic understanding), and masks high-entropy-shift positions to keep the low tail. This improves out-of-distribution generalization, and the resulting judges outperform outcome-supervised RL by 2-9 percentage points on subjective subcategories while staying competitive on objective ones.
More from Research
- ARC Prize finds Qwen3.8-27B's chat template injects different instructions per reasoning effort — GregKamradt · 2026-10-02
- DataScalar: The Failed 90s Memory-Centric Architecture That Aged Well — lauriewired · 2026-10-02
- AAAI review reform targets 'micro result' papers of minor ablations, says Dietterich — tdietterich · 2026-10-02
- Eval builder explains anti-benchmaxxing: rotating all questions and constantly changing methodology — airesearch12 · 2026-10-02
- GOLLuM finds 36.3% of top-5% outcomes vs 29.7% for descriptor-based optimization — pschwllr · 2026-10-02
- Language model representations beat descriptors in multi-objective reaction optimization — pschwllr · 2026-10-02