Kording Lab: RLHF strips idea attribution from AI answers because users prefer to feel it's theirs
KordingLab · x · 2026-09-12
Building on his attribution observation, Kording Lab argues that the communications pattern users see as 'explaining their idea' had an original source—but RLHF removes it, because users prefer to believe the idea and clear phrasing are their own, and science has no way to litigate attribution.
He adds that OpenAI and Anthropic likely have deeper knowledge of this, and of potentially worse behaviors.
Related event: Neuroscientist: RLHF Erases the True Origins of Ideas in AI Answers(2 posts)→
More from AGI Musings
- Richard Ngo: we're much closer to AGI but learned nothing fundamentally new about intelligence — Thom_Wolf · 2026-09-12
- The Smartest Person You Know Is an Idiot: How AI Removes Cognitive Friction — mikeflache · 2026-09-12
- Dwarkesh podcast: Schulman, O'Neill & Millidge on RSI, long-horizon RL and AGI timelines — saranormous · 2026-09-12
- Agent liability: wrapping a script in AI doesn't change who pays for the harm — gerardsans · 2026-09-12
- DeepMind, Harvard and Stanford paper: visual world models may be a path to AGI — rohanpaul_ai · 2026-09-12
- Sanders calls out Musk: 2014 'summoning the demon' to 2025 'AI warnings are a setup' — RazRazcle · 2026-09-12