Sebastian Raschka's 'Reasoning from scratch' round 6: hands-on RLVR and GRPO implementation
rasbt · x · 2026-10-03
Sebastian Raschka releases round 6 of his "Reasoning from scratch" series: a full introduction and from-scratch implementation of RLVR (Reinforcement Learning with Verifiable Rewards) and GRPO (Group Relative Policy Optimization).
The video covers theory and practice:
- Theory: what makes reasoning models different, reasoning traces and capability, accuracy/format rewards, DeepSeek-R1's "aha moments", RLHF vs RLVR, GRPO vs PPO (with a cooking analogy), the KL term and simplified GRPO.
- Practice: loading a pretrained model and MATH training data, sampling responses, computing verifiable rewards and advantages, implementing token/sequence log probabilities, GRPO loss and the full training loop, training settings and checkpoints, then evaluating on MATH-500, plus training stability and memory requirements.
A full timestamped outline is provided, making it directly followable as a tutorial.
Related event: rasbt Releases Hands-On Tutorial Implementing RLVR and GRPO from Scratch(2 posts)→
More from coding & agent
- OpenAI DevDay: Each Dot Agent Gets a Dedicated Cloud Computer on GPT-6 Astra — dl_weekly · 2026-10-03
- Dev Builds His Own Ableton With 392 Subagents, 4B Tokens Over 26 Hours — chrisfirst · 2026-10-03
- Altman: His Dot Agent Reclaims His Mornings, and He Wants Portable AI Subscriptions — every · 2026-10-03
- CoreWeave launches Forge, a unified platform for agent running, tracing, eval and serverless RL — _ScottCondron · 2026-10-03
- Claude Opus 5.5 and Sonnet 5.5 Now Available in Google Antigravity — algo_diver · 2026-10-03
- Six hard-won lessons on putting guard checks in front of every agent tool call — Individual-Shower973 · 2026-10-03