Stanford CS 329H opener: 91% prefer vivid answers, exposing reward model bias
sanmikoyejo · x · 2026-10-08
Stanford's new course CS 329H: Machine Learning from Human Preferences (lecturer sanmikoyejo) opened with telling polls:
- Asked how to answer a kid's "why is the sky blue," 91% of students picked the vivid answer over the scientific one; on quicksort complexity, 88% picked the caveated answer—"a better answer" depends on who's asking.
- Implication: a reward model trained on pooled comparisons must average over that context, and every thumbs-up and arena vote carries the same problem.
- The course unifies tools from psychometrics, discrete choice, ML, decision theory, behavioral economics, and social choice—fields that rarely cite each other—teaching RLHF practitioners which assumptions their preference models make and when they fail.
More from Research
- NumanThabit aims to replace animal studies with sufficient physics simulation — MarwaEldiwiny · 2026-10-08
- Manuel Blum responds on matrix multiplication, citing his four 1964 conjectures — aran_nayebi · 2026-10-08
- Researcher's agent workflow: label 100 examples, have the agent scale to 10k pseudo-labels — ducha_aiki · 2026-10-08
- Biologists' most impactful move now: making their work verifiable for AI — neuroecology · 2026-10-08
- River's open recipe beats GPT-6 Astra Pro and Claude Opus 5.5 on text-to-SQL at <1% cost — ibab · 2026-10-08
- Humans and LLMs share working-memory limits: COLM paper points to one computational trade-off — SaxeLab · 2026-10-08