Small Model Post-Training Increments Portable to Large Models

tuvllms · x · 2026-07-09

This post introduces a new paper demonstrating that post-training parameter deltas from smaller models can be directly grafted onto larger models without needing to retrain the larger model from scratch.

The author notes that in certain settings, these grafted large models can even outperform their smaller "tutor" models, exhibiting a form of weak-to-strong generalization achieved during inference.

Original post →

More from Research

Research channel →