From RLHF to robot shaping: how human judgments become reward signals

binarybits · x · 2026-10-11

In a discussion with @andrewprock, binarybits draws a parallel between RLHF and robot shaping: in Christiano et al., human judgments are the data from which the machine infers a reward function — the open question being whether an agent can learn what humans mean by "better" from occasional comparisons, then generate its own dense training signal.

In robot shaping, the trainer's evaluations are effectively part of the reward-generation mechanism itself, turning the engineering question into how to reward, punish, decompose and structure training so the desired behavior emerges.

Original post →

More from AGI Musings

AGI Musings channel →