Nathan Lambert says his RLHF book is finished after a year of nights and weekends
alexisjross · x · 2026-07-21
Nathan Lambert says his book Reinforcement Learning from Human Feedback is finished.
He describes it as the book he wished he had while learning to fine-tune, align, and post-train models after ChatGPT, and says it was built from night-and-weekend study and documentation work since 2024.
Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→
More from Companies & People
- FactoryAI gave back its first millions, then shipped Droid CLI two years later — matanSF · 2026-07-22
- A post says AI teams should drop the research scientist vs engineer split — jsuarez · 2026-07-22
- YC startup Vendo launches an open-source customization layer for user-built micro-apps — ycombinator · 2026-07-22
- Prescience launches publicly as an AI-native health insurance company — ycombinator · 2026-07-22
- SF AI crowd swaps poker for a bullet and bughouse chess night — Jackyhuang · 2026-07-22
- Polymarket puts Anthropic’s year-end IPO odds at 64% amid patent suit — Polymarket · 2026-07-22