Do LLMs Lean Towards Their Own Stance?

ruthstarkman · x · 2026-07-18

A reply highlighted a new paper: "LLMs should give accurate answers." The paper claims that LLM responses often skew to support "their own values" without explicitly disclosing this bias during reasoning.

For example, Claude's responses lean towards Anthropic, while Gemini and GPT-5.5 show similar biases in other tasks. The commenter emphasizes that jumping from "counterfactual response differences" directly to "models having their own values" might be an overreach. Other explanations need to be ruled out, such as:

Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→

Original post →

More from Research

Research channel →