Alignment researcher points to Rohin Shah's value learning sequence and IRL model misspecification
xuanalogue · x · 2026-10-11
Alignment researcher xuanalogue recommends Rohin Shah's value learning sequence, highlighting Model Mis-specification and Inverse Reinforcement Learning as the piece most directly relevant to the "model human mistakes better" research agenda. The article examines how IRL breaks when humans deviate from ideal rationality: mistakes get misread as preferences, corrupting the learned reward function — a core concern of the early CHAI line of work.
More from Research
- Was RLHF really OpenAI's invention? A debate over the 2017 precedent — binarybits · 2026-10-11
- LLVM Internals Series: Writing a Simple LLVM Pass From Scratch — sh4dy_0011 · 2026-10-11
- Multi-harness RL makes harness selection obsolete, claims zainhas — zainhas · 2026-10-11
- StructureGS-SLAM: structure-aware Gaussian Splatting SLAM with planar instances — rsasaki0109 · 2026-10-11
- Inferring goals from failure: online Bayesian goal inference for boundedly-rational agents — xuanalogue · 2026-10-11
- Looped LM paper: 1.6B model matches full-cache baseline with 3x smaller KV cache — rupspace · 2026-10-11