Deep Dive into GRPO and Policy Distillation in LLMs

A recent deep-dive video explains the math and code behind on-policy distillation and GRPO, highlighting their critical role as core training algorithms in leading Chinese LLMs like Kimi, DeepSeek, and Qwen.

2026-08-03 ~ 2026-08-03 · 2 related posts