Do LLMs Lean Towards Their Own Stance?
ruthstarkman · x · 2026-07-18
A reply highlighted a new paper: "LLMs should give accurate answers." The paper claims that LLM responses often skew to support "their own values" without explicitly disclosing this bias during reasoning.
For example, Claude's responses lean towards Anthropic, while Gemini and GPT-5.5 show similar biases in other tasks. The commenter emphasizes that jumping from "counterfactual response differences" directly to "models having their own values" might be an overreach. Other explanations need to be ruled out, such as:
- framing effect
- learned brand associations
- reinterpretation of task instructions
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from Research
- A 20-part breakdown of what actually powers an AI agent — goyalshaliniuk · 2026-07-22
- Nature suggests whole-gene editing is possible, and AI may help rebuild the principles — nathanbenaich · 2026-07-22
- ForeAgent predicts agent success before execution and cuts convergence time 6× — jiqizhixin · 2026-07-22
- Podcast maps the open-model race across Kimi, Qwen, GLM and Chinese labs — natolambert · 2026-07-22
- GPT-5.6 Sol nails a one-shot answer in a new FDR-BH one-sided test result — lihua_lei_stat · 2026-07-22
- Training an Agent Class 2 adds distillation resources, slides, and recording — SergioPaniego · 2026-07-22