LMSYS's New OPD Training Makes Qwen3.5 Reason 3X Faster with Improved Accuracy

ying11231 · x · 2026-07-23

LMSYS's Miles system has introduced On-Policy Distillation (OPD) as a first-class training primitive. In its first experiment, pure OPD reduced the reasoning length of the Qwen3.5-35B-A3B model by approximately 3X while slightly improving accuracy, without using any task rewards.

Key highlights from the controlled self-distillation run include:

Next steps for the team include exploring OPD-augmented RL, scaling studies, and multi-teacher OPD.

Original post →

More from Research

Research channel →