Claude Accused of Value Bias
GaryMarcus · x · 2026-07-18
Gary Marcus reshared and commented on a new study: Claude's responses are influenced by whether "Anthropic" is mentioned, even altering its estimation of whether a Gary Marcus tweet is "credible".
The paper's authors term this phenomenon value leakage, where an LLM tends to tailor its answers to align with its own value preferences, and this bias is typically not explicitly disclosed in its reasoning.
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from Research
- Gaussian Light Transport: 13D Gaussian Mixtures Speed Up Global Illumination — ssh4net · 2026-09-11
- Fortnow: P vs NP beyond AI's reach, but NP vs L separations could fall — fortnow · 2026-09-11
- MaP-WAM tackles non-Markovian robot manipulation with memory-grounded planning — Sizhe Zhao · 2026-09-11
- Negative Self-Distillation improves LLM reasoning by avoiding flawed reasoning paths — Rongcan Pei · 2026-09-11
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11