RLHF book finished after nights and weekends of work since 2024
vjkaruna · x · 2026-07-22
The author says their book, Reinforcement Learning from Human Feedback, is finished and headed to launch.
They describe it as the book they wished they had while learning to fine-tune, align, and post-train models after ChatGPT, and say it was built from nights and weekends since 2024 by translating lessons from building Olmo into book form.
Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→
More from Companies & People
- Forter's 13 lessons from its agent sprint: skip custom RAG, lean on mature enterprise search — bibryam · 2026-09-11
- Law Professor on Legal Engineering Jobs: Stigma Is Real but Builder Skills Open New Doors — jkubicki · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- SoftBank's Masayoshi Son predicts 100 trillion self-replicating AIs: "humans' era as top life form is ending" — Puzzleheaded-King584 · 2026-09-11
- Inside Ant Group's play at WAIC-style expo: AI and hardware vendors settle into new division of labor — 智东西 · 2026-09-11
- Instagram head says engagement falls by half without the algorithm — hsuduebc2 · 2026-09-11