SAP's 1B Reranker: On-Policy Distillation + Off-Policy GRPO Beats Offline KD by 4.6 nDCG Points
_reachsumit · x · 2026-09-03
SAP researchers propose a two-stage RL-based framework for training compact instruction-following rerankers.
- Stage 1: A 4B teacher reranker is strengthened with off-policy GRPO using LLM-judge feedback on 88K instruction-following examples.
- Stage 2: A 1B student samples rankings from its own policy and receives soft teacher-derived rewards on them, coupling student exploration with knowledge transfer — unlike conventional offline imitation of teacher outputs on fixed examples.
- Results: The student hits 0.7670 nDCG@6 on MAIR-11 (11 subsets, 869 queries), +4.6 points over offline listwise KD. Controlled comparisons show neither pairwise RankNet KD nor on-policy GKD reproduces the gains, which persist across MAIR-Full (126 tasks, 9,356 queries), with the largest advantages under distribution shift.
More from Research
- Nora optimizer keeps Muon's matrix structure benefits without the full cost — burkov · 2026-09-03
- UCLA x Google unveil PaperBanana-Interact: multi-turn chat refinement for scientific diagrams — kaiwei_chang · 2026-09-03
- Unverified claim: 200-digit number posted that supposedly divides RSA-260 — marvinvonhagen · 2026-09-03
- HCI papers increasingly use LLM judges while obfuscating it, researcher warns — IanArawjo · 2026-09-03
- Why VRChat Particle Pools Still Work: Gravity Naturally Converges States — Michael_Moroz_ · 2026-09-03
- Radix Sort Hit 5B Key-Value Pairs per Second With Zero Compute Shaders — Michael_Moroz_ · 2026-09-03