OASIS fixes on-policy self-distillation's scale collapse, gaining 3+ points over OPSD at 8B

Md. Ismail Hossain · hf · 2026-10-01

This work diagnoses why on-policy self-distillation (OPSD) fails to scale for LLM reasoning.

Original post →

More from Research

Research channel →