Deep Dive: How GRPO and On-Policy Distillation Power Frontier LLMs

johnolafenwa · reddit · 2026-08-03

The author highlights that recent technical reports from frontier models like Kimi, DeepSeek, Qwen, and GLM heavily rely on on-policy distillation (OPD) and GRPO-style algorithms.

To explain these concepts, the author published a deep dive video detailing the mathematics and code behind these algorithms. The tutorial explores how they connect to pretraining and supervised fine-tuning (SFT), helping developers better understand the current LLM training paradigm.

Related event: Deep Dive into GRPO and Policy Distillation in LLMs(2 posts)→

Original post →

More from Research

Research channel →