CaRE-KD accepted at NeurIPS 2026: letting students decide when to trust distillation teachers
Tanmoy_Chak · x · 2026-10-03
- In knowledge distillation, teachers hallucinate, yet standard KD forces the student to copy them anyway, wasting probability mass on noisy low-confidence predictions and overwriting correct student priors.
- CaRE-KD (accepted at NeurIPS 2026) lets the student decide when to trust the teacher:
- Token level: a confidence gate switches between mean-seeking and mode-seeking divergence based on teacher confidence.
- Sequence level: a rejection mechanism drops updates where the teacher is epistemically less reliable than the student.
- Preprint available.
More from Research
- RSI Arena live-streams AI agents training models, judged by human evaluation — kuchaev · 2026-10-03
- BIND Action Head Ties Robot Actions to 2D Image Features for Data-Efficient Policies — yuewang314 · 2026-10-03
- JEPA-TTT: Persistent test-time training lifts world model planning by 153% — DanielKhashabi · 2026-10-03
- ByteDance's DMAD distills MiniMax H3 video to 4 steps — AgeNo5351 · 2026-10-03
- Honeycomb: constant-size HexMemory keeps video world models consistent over long-horizon generation — Jack Wei Lun Shi · 2026-10-03
- X-Tree tokenizes reusable experience into skill hierarchies, boosting agent RL by up to 5.8% — UWaterloo · 2026-10-03