New on-policy distillation method lets student models match or exceed their teachers across four pairs
_akhaliq · x · 2026-10-02
The paper "The Teacher Is a Direction, Not a Destination" introduces a new on-policy distillation method that extrapolates RL-induced representation residuals, letting students continue along the direction their teacher points to rather than stopping at the teacher's ceiling.
Key points:
- Students match or exceed their teachers across four model pairs
- Code and checkpoints are planned for release
This challenges the standard distillation assumption that students can only approach teacher performance.
More from Research
- Apple paper: structured selection-based reasoning cuts search agent latency by 90% — _reachsumit · 2026-10-02
- GrIS paper reframes Semantic IDs as recursive graph partitioning for generative recommendation — _reachsumit · 2026-10-02
- MatRAG pairs hierarchical clustering with Matryoshka embeddings to cut multi-hop RAG cost — _reachsumit · 2026-10-02
- Meta paper: only 50-60% of recommendation training time actually trained before optimizations — _reachsumit · 2026-10-02
- OmniSeek turns Omni-LLMs into agents that actively seek audio-visual evidence — Haibo Wang · 2026-10-02
- Netflix's Align Then Reason lip-sync judge boosts mean AUC by up to 59% — netflix · 2026-10-02