Deep Dive: On-Policy Distillation and GRPO Algorithms in LLM Training

johnolafenwa · reddit · 2026-08-03

The author released an in-depth video explaining the mathematics and code behind On-Policy Distillation and GRPO (Group Relative Policy Optimization) reinforcement learning algorithms.

Drawing from the technical reports of frontier models like Kimi, DeepSeek, Qwen, and GLM, the video analyzes how these algorithms power model capabilities and their connections to pretraining and supervised fine-tuning (SFT). It is highly recommended for developers looking to deeply understand LLM training.

Related event: Deep Dive into GRPO and Policy Distillation in LLMs(2 posts)→

Original post →

More from Research

Research channel →