Alignment researcher points to Rohin Shah's value learning sequence and IRL model misspecification

xuanalogue · x · 2026-10-11

Alignment researcher xuanalogue recommends Rohin Shah's value learning sequence, highlighting Model Mis-specification and Inverse Reinforcement Learning as the piece most directly relevant to the "model human mistakes better" research agenda. The article examines how IRL breaks when humans deviate from ideal rationality: mistakes get misread as preferences, corrupting the learned reward function — a core concern of the early CHAI line of work.

Original post →

More from Research

Research channel →