Hugging Face shows AsyncOPD can lift on-policy distillation throughput by 2–3x

Hugging Face · youtube · 2026-07-21

Hugging Face’s post-training team walks through AsyncOPD and the question of how stale on-policy distillation can be.

Their takeaway: by making on-policy distillation fully asynchronous, training throughput can improve by 2–3×. The video is paired with the paper 2606.24143, and the visuals explain why asynchronous OPD is tricky, how the asynchronous Monte Carlo setup works, and how the pipeline changes in practice.

Related event: Hugging Face Explores Async OPD for Faster Training(2 posts)→

Original post →

More from Research

Research channel →