SAP's 1B Reranker: On-Policy Distillation + Off-Policy GRPO Beats Offline KD by 4.6 nDCG Points

_reachsumit · x · 2026-09-03

SAP researchers propose a two-stage RL-based framework for training compact instruction-following rerankers.

Original post →

More from Research

Research channel →