Nathan Lambert publishes a full RLHF and post-training course for language models

natolambert · x · 2026-07-29

Nathan Lambert’s RLHF & Post-Training course collects a full set of materials on reinforcement learning from human feedback and language-model post-training.

The course includes lecture videos, slides, PDFs, source materials, prerequisites, and extra resources. It is aimed at early AI graduate students but is designed to be accessible to motivated learners who can use modern LLMs as tutors while working through the math, code, and jargon.

Original post →

More from Research

Research channel →