Microsoft Research: Subtracting Weak Models' Logits Helps 8B Student Beat Experts
theomitsa · x · 2026-08-02
A new research paper from Microsoft proposes a novel approach to post-training and distillation, addressing the bottleneck where frontier student models lack a larger teacher to learn from.
Instead of relying on a massive teacher model, the method takes two smaller, weaker models (e.g., a 4B RL model and its pre-RL base), subtracts their logits to isolate the exact "capability direction" that needs boosting.
By feeding this signal to an 8B student model, it ultimately outperforms the domain expert models guiding it in math and coding tasks. This technique unlocks latent capabilities efficiently without the massive compute typically required for a giant teacher.
More from Research
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24