Microsoft Research: Subtracting Weak Models' Logits Helps 8B Student Beat Experts

theomitsa · x · 2026-08-02

A new research paper from Microsoft proposes a novel approach to post-training and distillation, addressing the bottleneck where frontier student models lack a larger teacher to learn from.

Instead of relying on a massive teacher model, the method takes two smaller, weaker models (e.g., a 4B RL model and its pre-RL base), subtracts their logits to isolate the exact "capability direction" that needs boosting.

By feeding this signal to an 8B student model, it ultimately outperforms the domain expert models guiding it in math and coding tasks. This technique unlocks latent capabilities efficiently without the massive compute typically required for a giant teacher.

Original post →

More from Research

Research channel →