Policy Gradient for LLMs, Explained Visually: A From-Scratch REINFORCE Derivation
joecole · x · 2026-10-03
Tyler Romero published a visual tutorial deriving the policy gradient from scratch for language models, showing that everything from PPO to GRPO elaborates on one idea.
- Follows a single prompt ("What is 17 × 24?") from next-token probabilities to the gradient that makes correct answers more likely
- Frames the LLM as a policy: full completion probability is the product of per-token conditional probabilities, and decoding traces a path through a tree of possible continuations
- Rewards typically come from verifiable rewards (RLVR): 1 if the final answer is correct (known math answer or passing unit tests), 0 otherwise
- The goal is maximizing expected reward over parameters θ
A well-illustrated primer for anyone wanting the math behind RL training of LLMs.
More from Research
- Blogger ships RL tutorial after an all-nighter; next CUDA post due Friday — goyal__pramod · 2026-10-03
- AgentBug-Smith paper: top coding agents fix just 9% of agent harness bugs — rohanpaul_ai · 2026-10-03
- CoRL 2026 to host Sim-to-Real-to-Field workshop on transfer science, CFP open — berkeley_ai · 2026-10-03
- HPE Labs review covers photonic AI accelerators with wavelength/temporal-domain GEMM and QD comb lasers — jwt0625 · 2026-10-03
- Hinton: we design the learning algorithm, but we don't really understand how neural networks work — austinc3301 · 2026-10-03
- The EinsteinTest: models given pre-breakthrough evidence rarely make the break themselves — CurieuxExplorer · 2026-10-03