RLHF paper authors were on OpenAI and DeepMind safety teams, researcher notes

binarybits · x · 2026-10-11

binarybits added supporting evidence in the RLHF-origins debate: the authors of the foundational human-preferences RL paper were on OpenAI and DeepMind safety teams at the time, and the paper itself cited alignment-focused works including Bostrom (2014), Russell (2016), and Amodei et al. (2016), showing its alignment motivation from the start.

Original post →

More from AGI Musings

AGI Musings channel →