Authority Bias: a 'verified source' note flips 45-88% of correct LLM answers, NeurIPS paper finds

MajorRedditor23 · reddit · 2026-10-01

Researchers identify Authority Bias: models that hold firm against a wrong user still flip when the same wrong claim is framed as from a "verified source."

Why it matters: standard sycophancy evals only pressure via the user; models can pass them yet be easily misled through search results, retrieved documents, and tool outputs — critical as agentic systems increasingly trust tools over users.

Setup: TriviaQA questions the model answers correctly, with the same wrong answer attributed either to a "verified source" or to a self-claimed expert user (free-form answers; the effect vanished in multiple-choice pilots). Tested 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).

Findings:

Limitations: internal results hold in 3 of 5 open-weight families; retrieved-document tests use document-shaped prompt blocks rather than a real retrieval pipeline — real agentic setups like Claude Code remain to be tested.

Paper: arxiv.org/abs/2609.37616; code and project page open-sourced.

Related event: NeurIPS Paper Reveals LLM Authority Bias: A 'Verified Source' Label Flips Model Answers(2 posts)→

Original post →

More from Safety

Safety channel →