Nat Lambert finishes a RLHF book built from nights and weekends since 2024
danielhanchen · x · 2026-07-22
Nat Lambert says his book Reinforcement Learning from Human Feedback is finished, describing it as the guide he wished he had when learning to fine-tune, align, and post-train models after ChatGPT. He says the book was assembled through nights-and-weekends study since 2024, drawing on lessons from building Olmo.
Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→
More from Companies & People
- Forter's 13 lessons from its agent sprint: skip custom RAG, lean on mature enterprise search — bibryam · 2026-09-11
- Law Professor on Legal Engineering Jobs: Stigma Is Real but Builder Skills Open New Doors — jkubicki · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- SoftBank's Masayoshi Son predicts 100 trillion self-replicating AIs: "humans' era as top life form is ending" — Puzzleheaded-King584 · 2026-09-11
- Inside Ant Group's play at WAIC-style expo: AI and hardware vendors settle into new division of labor — 智东西 · 2026-09-11
- Instagram head says engagement falls by half without the algorithm — hsuduebc2 · 2026-09-11