PUST: New Paradigm Uses Small Models to Guide Large Model Training

A new post-training paradigm named PUST decouples exploration from distribution alignment by using small proxy models to explore and transfer update signals, reducing the high costs of reinforcement learning for large models.

2026-07-14 ~ 2026-07-15 · 2 related posts