Post-Training AI book posts draft chapters: SFT and GRPO in a few hundred lines
ben_burtenshaw · x · 2026-08-25
The author shared the first pre-release chapters of Post-Training AI: A Practical Guide to Fine-Tuning and Reinforcement Learning. The book walks a simple line through post-training to build intuition for how agents learn across SFT, GRPO, distillation, and environments, deliberately skipping some concepts in favor of clarity.
Chapter 1 defines post-training — what each stage changes, what none of them can fix, and why the techniques now work outside frontier labs. Chapter 2 is a speed run implementing SFT and GRPO in a few hundred lines of code.
More from Research
- Visualizing α-geodesics on the probability simplex: exponential vs mixture geometry — FrnkNlsn · 2026-08-25
- IBM Open-Sources Granite-4.2-30B: Built-in Chain-of-Thought, 512K Context, Apache 2.0 — jacek2023 · 2026-08-25
- Sliding puzzle video explains why AI reasoning needs guided search, not just sampling — CatAstro_Piyush · 2026-08-25
- Anthropic and AWS sponsor hackathon to decode rare disease using open genome data — CatAstro_Piyush · 2026-08-25
- Contravariance theory for RSA: removing nuisance components restores full contravariance — dyamins · 2026-08-25
- Papers with Code Updates Healthcare AI Page with Benchmarks and Papers — NielsRogge · 2026-08-25