Mid-Training Distillation Boosts Reasoning But Small Models Lose Fact Recall First
roydanroy · x · 2026-09-02
- Citing a KD study: forward KL distillation from post-trained teachers behaves fundamentally differently during mid-training, improving reasoning and factual recall relative to standard next-token prediction.
- The reply links arXiv 2310.04680 (The Cost of Down-Scaling Language Models): via both pruning and dense scaling, shrinking a model by >30% sharply degrades pre-training fact recall, while a 60–70% reduction largely preserves in-context learning.
- Together the threads show compression/distillation doesn't shrink capabilities uniformly — memorized facts are the most fragile.
More from Models
- Tencent open-sources WeMM-Embedding: unified text/image/video embeddings, Apache 2.0 — tomaarsen · 2026-09-02
- When nobody can track frontier model progress, closed-model business may lose to open weights — StewartalsopIII · 2026-09-02
- Looped transformer is no dark art: rasbt debunks the OpenAI Astra rumor — rasbt · 2026-09-02
- GLM 5.2 slug references spotted in Google Antigravity CLI, hinting at integration — gaganghotra_ · 2026-09-02
- Anthropic's Fable-5.1-max Grabs #4 on eyebench-v3 in Biggest Bench Jump Yet — adonis_singh · 2026-09-02
- Dev Tries Claude for Coding Agents Again, Finds Sol 5.6 Sharper But Less Friendly — iskander · 2026-09-02