Post-Training AI book posts draft chapters: SFT and GRPO in a few hundred lines

ben_burtenshaw · x · 2026-08-25

The author shared the first pre-release chapters of Post-Training AI: A Practical Guide to Fine-Tuning and Reinforcement Learning. The book walks a simple line through post-training to build intuition for how agents learn across SFT, GRPO, distillation, and environments, deliberately skipping some concepts in favor of clarity.

Chapter 1 defines post-training — what each stage changes, what none of them can fix, and why the techniques now work outside frontier labs. Chapter 2 is a speed run implementing SFT and GRPO in a few hundred lines of code.

Original post →

More from Research

Research channel →