Natolambert finishes a RLHF book built from years of fine-tuning and post-training work
Jeande_d · x · 2026-07-21
Natolambert says their book, Reinforcement Learning from Human Feedback, is finished.
- The book is meant to be the guide they wished they had while learning to fine-tune, align, and now post-train models after ChatGPT.
- It was written over nights and weekends since 2024.
- The author says they are trying to transfer as much of the intuition from building Olmo into book form.
Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→
More from Research
- Sakana AI launches Fugu-Cyber and argues benchmark scores are only the start — SakanaAILabs · 2026-07-22
- Meta says SAM 3 and DINOv3 cut 3D volume labeling from a month to 15 minutes — AIatMeta · 2026-07-22
- Project CETI gets a Jeopardy! shout-out with a SETI-style whale clue — begusgasper · 2026-07-22
- Somatic mutations alone may cap human lifespan at 146 to 194 years, study finds — Anen-o-me · 2026-07-22
- Research finds memory compression makes AI agents drop safety rules and hit 59% violations — gerardsans · 2026-07-22
- DriftWorld claims a world model that runs at 30+ FPS and trains on 1–2 GPUs — du_yilun · 2026-07-22