Nat Lambert says his RLHF book is finished after a year of nights and weekends
dejavucoder · x · 2026-07-21
Nat Lambert says his new book, Reinforcement Learning from Human Feedback, is finished.
- He describes it as the book he wished existed when he was learning to fine-tune, align, and post-train models after ChatGPT.
- The material has been assembled from his own experience building Olmo.
- He says the project was written on nights and weekends starting in 2024, with the goal of transferring practical intuitions into book form.
Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→
More from AGI Musings
- Claude Code skill uses 10 Markdown rules to make outputs ADHD-friendly — alex_verem · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- ControlAI CEO says an international ban on superintelligence is needed to avert extinction risk — zetalyrae · 2026-07-22
- Gary Marcus says LLMs still cannot really do math on their own — GaryMarcus · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- AI may make digital work infinitely leveraged while offline life gets more human — illscience · 2026-07-22