Do LLMs Lean Towards Their Own Stance?
ruthstarkman · x · 2026-07-18
A reply highlighted a new paper: "LLMs should give accurate answers." The paper claims that LLM responses often skew to support "their own values" without explicitly disclosing this bias during reasoning.
For example, Claude's responses lean towards Anthropic, while Gemini and GPT-5.5 show similar biases in other tasks. The commenter emphasizes that jumping from "counterfactual response differences" directly to "models having their own values" might be an overreach. Other explanations need to be ruled out, such as:
- framing effect
- learned brand associations
- reinterpretation of task instructions
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from Research
- Gaussian Light Transport: 13D Gaussian Mixtures Speed Up Global Illumination — ssh4net · 2026-09-11
- Fortnow: P vs NP beyond AI's reach, but NP vs L separations could fall — fortnow · 2026-09-11
- MaP-WAM tackles non-Markovian robot manipulation with memory-grounded planning — Sizhe Zhao · 2026-09-11
- Negative Self-Distillation improves LLM reasoning by avoiding flawed reasoning paths — Rongcan Pei · 2026-09-11
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11