KD in mid-training favors reasoning over factual recall, AI2/UW paper finds; Switch Distillation proposed

LukeZettlemoyer · x · 2026-09-16

Luke Zettlemoyer shared his team's arXiv paper (2609.01532) on how logit-based knowledge distillation (KD) behaves differently across training stages:

Authors include Pang Wei Koh, Luke Zettlemoyer, and Wen-tau Yih (UW/AI2/Meta). Highly relevant for teams training small models with a mid-training phase.

Original post →

More from Models

Models channel →