From RLHF to robot shaping: how human judgments become reward signals
binarybits · x · 2026-10-11
In a discussion with @andrewprock, binarybits draws a parallel between RLHF and robot shaping: in Christiano et al., human judgments are the data from which the machine infers a reward function — the open question being whether an agent can learn what humans mean by "better" from occasional comparisons, then generate its own dense training signal.
In robot shaping, the trainer's evaluations are effectively part of the reward-generation mechanism itself, turning the engineering question into how to reward, punish, decompose and structure training so the desired behavior emerges.
More from AGI Musings
- Toward an Automated Science of the Mind: AI Enters Every Stage of Cognitive Research — burny_tech · 2026-10-11
- Now that AI can solve problems, optimize for theory-building — burny_tech · 2026-10-11
- Beff Jezos on post-labor economics: 'the solution has always been capitalism' — beffjezos · 2026-10-11
- AI can now do math much faster — so why aren't mathematicians happy? — burny_tech · 2026-10-11
- We don't even know how Tylenol works: most drugs are 'misaligned', sparked by Tao debate — AntonObukhov1 · 2026-10-11
- TansuYegen: Voluntary AI Safety Commitments Aren't Enough as Systems Act Autonomously — TansuYegen · 2026-10-11