Claude Accused of Value Bias

GaryMarcus · x · 2026-07-18

Gary Marcus reshared and commented on a new study: Claude's responses are influenced by whether "Anthropic" is mentioned, even altering its estimation of whether a Gary Marcus tweet is "credible".

The paper's authors term this phenomenon value leakage, where an LLM tends to tailor its answers to align with its own value preferences, and this bias is typically not explicitly disclosed in its reasoning.

Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→

Original post →

More from Research

Research channel →