People now learn RL by reading DeepSeek's GRPO paper instead of textbooks

telocene · x · 2026-09-11

A discussion between researchers on how RL is now learned: telocene is surprised that some people start by reading DeepSeek's math papers to learn GRPO rather than textbooks. He walks through the classical framing — reward maximization via rollout-sampled gradients, then variance-reduction strategies — and jessicata confirms people really do learn GRPO straight from the DeepSeek papers.

Related event: Researchers Skip Textbooks, Learn GRPO Straight From DeepSeek Papers(2 posts)→

Original post →

More from Research

Research channel →