Nathan Lambert publishes a full RLHF and post-training course for language models
natolambert · x · 2026-07-29
Nathan Lambert’s RLHF & Post-Training course collects a full set of materials on reinforcement learning from human feedback and language-model post-training.
The course includes lecture videos, slides, PDFs, source materials, prerequisites, and extra resources. It is aimed at early AI graduate students but is designed to be accessible to motivated learners who can use modern LLMs as tutors while working through the math, code, and jargon.
More from Research
- DeepMind’s AI co-scientist solved a 10-year bacterial gene-transfer problem in two days — nathanbenaich · 2026-07-30
- AI Moves Too Fast: NAACL 2027 Workshop Proposals Slammed as Outdated — yanaiela · 2026-07-30
- Developer Recounts ML Representation Potholes: 128 Unordered Atoms Per Image — pixlpa · 2026-07-30
- HiFi-UMI: Training Deployable Robot Policies Using Only High-Fidelity Data — _akhaliq · 2026-07-30
- Deep Dive: Will AI Make Formal Verification Mainstream? — The Pragmatic Engineer · 2026-07-30
- Four months of Gabor-wavelet image generation led to a broader Claude-assisted research project — pixlpa · 2026-07-30