RIDE: Extrapolating RL-Induced Representation Residuals Lets Distilled Students Match or Beat RL Teachers
Hao Li · hf · 2026-10-01
- Problem: Output-space extrapolation in on-policy distillation is unstable because the LM head attenuates hidden-state changes anisotropically and log-prob ratios inject amplified noise.
- Key observation: RL shifts a model's internal representations relative to its base checkpoint, and this shift direction is measurable at every layer.
- Method (RIDE): Extrapolate the RL-induced change directly in representation space — at each layer and token position, compute the residual between the teacher and its pre-RL checkpoint, then regress the student's hidden states toward targets displaced beyond the teacher along that residual. This is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher.
- Results: Across four base/RL-teacher pairs spanning scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL teacher on every pair (the only method to do so on average) and consistently outperforms output-space extrapolation, which degrades students when the teacher is near its base.
More from Research
- ArchMap lands in Nature Genetics: code-free single-cell mapping onto reference atlases — burny_tech · 2026-10-01
- New Paper Argues Literary Tools Are Essential for Building Culturally Literate AI — begusgasper · 2026-10-01
- NVIDIA's Instant NuRec reconstructs a drivable 3DGS world from driving logs in ~1.5 seconds — rsasaki0109 · 2026-10-01
- KV-streams trains SWE agents 2x faster by preserving KV cache across compaction — burny_tech · 2026-10-01
- Researcher argues "ego" beats "persona" for describing LLM identity — repligate · 2026-10-01
- Silicon microring modulators push past 200Gb/s per lane to cut AI optical I/O power — jwt0625 · 2026-10-01