Was RLHF really OpenAI's invention? A debate over the 2017 precedent

binarybits · x · 2026-10-11

binarybits and andrewprock argue over RLHF's historical credit. binarybits frames the key insight as "train a reward model from pairwise human feedback, then use it for RL on another model," asking whether any pre-2017 paper used this exact technique. andrewprock counters that using human feedback is old news in ML, was not invented by OpenAI, and calling it comparable to inventing the lightbulb is "facile poppycock," citing an early Robot Shaping paper. The crux: is the novelty the general idea of human feedback, or the specific pairwise-preference-to-reward-model-to-RL pipeline?

Related event: Debate Erupts Over Who Really Invented RLHF(2 posts)→

Original post →

More from Research

Research channel →