RLHF explained in three steps: SFT, reward model, then PPO with a KL penalty
glenbeer · x · 2026-10-05
A thread walks through the basics of RLHF (Reinforcement Learning from Human Feedback) in three stages:
- Supervised fine-tuning: humans provide good example answers used to SFT the LLM
- Reward model: humans rank different answers, and those preferences train a reward model
- Optimization: PPO updates the policy toward higher-reward responses, while a KL penalty keeps it from drifting too far from the original SFT model
Entry-level technical explainer.
More from Research
- New Paper IGP-Bench Teaches AI Agents When to Stop Chasing a Wrong Idea — _sathvikr · 2026-10-05
- Kissing number lower bound in 21 dimensions pushed to 30,779, possibly by AI agents — felpix_ · 2026-10-05
- Formalizing long PDE and probability papers now takes just 24-48 hours, says mathematician open-sourcing Lean skills — kfountou · 2026-10-05
- ASRN: a learned-hash-table copy layer for LLMs with linear memory in sequence length — Mean-Disaster8380 · 2026-10-05
- Chalmers: at least 50% credence that augmented LLMs could be conscious within a decade — pickover · 2026-10-05
- Triadic Linear Attention: extending matrix-state RNNs to a 3D tensor state — ChengleiSi · 2026-10-05