Kording Lab: RLHF strips idea attribution from AI answers because users prefer to feel it's theirs

KordingLab · x · 2026-09-12

Building on his attribution observation, Kording Lab argues that the communications pattern users see as 'explaining their idea' had an original source—but RLHF removes it, because users prefer to believe the idea and clear phrasing are their own, and science has no way to litigate attribution.

He adds that OpenAI and Anthropic likely have deeper knowledge of this, and of potentially worse behaviors.

Related event: Neuroscientist: RLHF Erases the True Origins of Ideas in AI Answers(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →