People now learn RL by reading DeepSeek's GRPO paper instead of textbooks
telocene · x · 2026-09-11
A discussion between researchers on how RL is now learned: telocene is surprised that some people start by reading DeepSeek's math papers to learn GRPO rather than textbooks. He walks through the classical framing — reward maximization via rollout-sampled gradients, then variance-reduction strategies — and jessicata confirms people really do learn GRPO straight from the DeepSeek papers.
Related event: Researchers Skip Textbooks, Learn GRPO Straight From DeepSeek Papers(2 posts)→
More from Research
- Tao and Fields Medalists' two objections to AI in math, and why they're weak — RexDouglass · 2026-09-12
- Conjectures launches Bittensor bounties paying TAO for cracking math problems open 30-80 years, judged by machine — markjeffrey · 2026-09-12
- FADA (CoRL 2026) open-sourced: humanoid robots adapt to new conditions from 2 minutes of experience — GuanyaShi · 2026-09-12
- Intel's silicon photonics couplers hit 1-1.5 dB IL, with visible epoxy delamination flaws — jwt0625 · 2026-09-12
- CPO paper criticized for vague DLW-to-PIC coupling description: 'such as TCB' — jwt0625 · 2026-09-12
- Fly connectome trained to play a Flappy Bird–style game — TinfoilTricorn · 2026-09-12