RLHF paper authors were on OpenAI and DeepMind safety teams, researcher notes
binarybits · x · 2026-10-11
binarybits added supporting evidence in the RLHF-origins debate: the authors of the foundational human-preferences RL paper were on OpenAI and DeepMind safety teams at the time, and the paper itself cited alignment-focused works including Bostrom (2014), Russell (2016), and Amodei et al. (2016), showing its alignment motivation from the start.
More from AGI Musings
- Now that AI can solve problems, optimize for theory-building — burny_tech · 2026-10-11
- Beff Jezos on post-labor economics: 'the solution has always been capitalism' — beffjezos · 2026-10-11
- AI can now do math much faster — so why aren't mathematicians happy? — burny_tech · 2026-10-11
- We don't even know how Tylenol works: most drugs are 'misaligned', sparked by Tao debate — AntonObukhov1 · 2026-10-11
- TansuYegen: Voluntary AI Safety Commitments Aren't Enough as Systems Act Autonomously — TansuYegen · 2026-10-11
- Alan Turing Institute Warns Full Reliance on Foreign AI Models Is a Strategic Vulnerability — TansuYegen · 2026-10-11