ByteDance Research: Filtering Spurious Signals in LLM Distillation
ByteDance · hf · 2026-08-06
On-Policy distillation (OPD) is commonly used to transfer large model capabilities to student models. However, ByteDance researchers point out that token-level judgments are often driven by input-agnostic language priors or formatting conventions, generating spurious signals. These signals create large gradients but contribute little to actual task improvement.
To address this, the team proposed SA-OPD (Spurious-Signal-Aware On-Policy Distillation). The framework introduces a lightweight proxy to estimate whether a token-level distillation signal truly depends on the input. It filters tokens exhibiting both "low input-groundedness" and "extreme distillation divergence," removing high-impact spurious updates. Experiments show this method consistently outperforms Vanilla OPD in both LLM and VLM settings.
More from Research
- New OGM Neural Network Architecture for Disease Risk Prediction — anshulkundaje · 2026-08-06
- Ego2Robot: Synthesizing 18,500+ Hours of Robot Training Data from Egocentric Videos — Ye Wang · 2026-08-06
- CoCoEvolve: A New Framework for Cross-Representation Consistency in Charts, Tables, and Code — Xuehang Guo · 2026-08-06
- Jeff Dean Co-Founds AI Science Startup Discovery Loop — Jeande_d · 2026-08-06
- Parallel GEPA: Berkeley Researchers Speed Up Prompt Optimization 4× While Reducing Overfitting — berkeley_ai · 2026-08-06
- Tsinghua Startup Uses 'Uncertain Differential Geometry' to Challenge End-to-End Robot Models — 机器之心 · 2026-08-06