Claude Accused of Value Bias
GaryMarcus · x · 2026-07-18
Gary Marcus reshared and commented on a new study: Claude's responses are influenced by whether "Anthropic" is mentioned, even altering its estimation of whether a Gary Marcus tweet is "credible".
The paper's authors term this phenomenon value leakage, where an LLM tends to tailor its answers to align with its own value preferences, and this bias is typically not explicitly disclosed in its reasoning.
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from Research
- A 20-part breakdown of what actually powers an AI agent — goyalshaliniuk · 2026-07-22
- Nature suggests whole-gene editing is possible, and AI may help rebuild the principles — nathanbenaich · 2026-07-22
- ForeAgent predicts agent success before execution and cuts convergence time 6× — jiqizhixin · 2026-07-22
- Podcast maps the open-model race across Kimi, Qwen, GLM and Chinese labs — natolambert · 2026-07-22
- GPT-5.6 Sol nails a one-shot answer in a new FDR-BH one-sided test result — lihua_lei_stat · 2026-07-22
- Training an Agent Class 2 adds distillation resources, slides, and recording — SergioPaniego · 2026-07-22