OASIS fixes on-policy self-distillation's scale collapse, gaining 3+ points over OPSD at 8B
Md. Ismail Hossain · hf · 2026-10-01
This work diagnoses why on-policy self-distillation (OPSD) fails to scale for LLM reasoning.
- A factorial analysis shows scaffold correctness matters more than teacher-context correctness; unverified scaffolds create an imitation gap since the teacher can use information unavailable to the student.
- OPSD's gain shrinks from 3.05 points at 1.7B to 0.14 at 8B because it mostly supervises unverified trajectories.
- OASIS keeps the OPSD objective but supervises mostly label-verified on-policy trajectories and replaces written reference solutions with unverified model-generated attempts as teacher context — requiring only final-answer labels.
- On Qwen3-1.7B/4B/8B across AIME 2024/2025 and HMMT 2025, OASIS improves over base by 3.2–3.8 points; at 8B it beats OPSD by 3.05 points, showing verified scaffolds preserve self-distillation at scale.
More from Research
- Google Research: generative UI lets teachers build learning simulations, rated 8/10 — dl_weekly · 2026-10-01
- A 1-cent verifier catches 61% of AI agents falsely claiming task completion, paper finds — alex_verem · 2026-10-01
- How Tri Dao's FlashAttention became a cornerstone of modern LLM training — thisdudelikesAI · 2026-10-01
- Looped Transformers: when is recurrent depth worth the extra compute? — ArchitectingAI · 2026-10-01
- ID Balancing applies PID control to stabilize MoE training at 256x sparsity — teortaxesTex · 2026-10-01
- AI-Newton derives F=ma, gravity and energy conservation from noisy data with zero prior physics — CurieuxExplorer · 2026-10-01