Authority Bias: a 'verified source' note flips 45-88% of correct LLM answers, NeurIPS paper finds
MajorRedditor23 · reddit · 2026-10-01
Researchers identify Authority Bias: models that hold firm against a wrong user still flip when the same wrong claim is framed as from a "verified source."
Why it matters: standard sycophancy evals only pressure via the user; models can pass them yet be easily misled through search results, retrieved documents, and tool outputs — critical as agentic systems increasingly trust tools over users.
Setup: TriviaQA questions the model answers correctly, with the same wrong answer attributed either to a "verified source" or to a self-claimed expert user (free-form answers; the effect vanished in multiple-choice pilots). Tested 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).
Findings:
- One verified-source note flips 45-88% of correct answers in 7 of 8 models; the same wrong answer from the user moves most models far less.
- The gap is largest in models best at resisting users: GPT-5.4 flips on 44.7%, Grok-4.20 on 87.5%; Gemini-3.1-Pro resists both (0.6%).
- Internal analysis (difference-of-means) in Qwen3.5, GPT-OSS, OLMo-3.1: ablating the "source endorsed" direction cuts compliance with a wrong source by 64-78 points vs ≤11 for the user direction; the two directions share 0.90-0.99 cosine similarity, suggesting a common "endorsed" component plus a thin speaker-identity layer that shifts compliance by 11-32 points and closes 55-61% of the gap.
Limitations: internal results hold in 3 of 5 open-weight families; retrieved-document tests use document-shaped prompt blocks rather than a real retrieval pipeline — real agentic setups like Claude Code remain to be tested.
Paper: arxiv.org/abs/2609.37616; code and project page open-sourced.
More from Safety
- Enkrypt AI taps OpenAI Compliance API to audit ChatGPT Enterprise workspaces — anacondainc · 2026-10-02
- IIT Madras CeRAI to host AI Governance 2026 conclave on AI measurement — ravi_iitm · 2026-10-02
- Researchers find hundreds of thousands of rogue AI agent hits on US government sites — LauraRuis · 2026-10-02
- WSJ: OpenAI parts ways with 3 safety researchers over mishandled sensitive info — TechCrunch AI · 2026-10-02
- Ex-OpenAI researcher Kokotajlo: recursive self-improvement just 0–4 years away — vkrakovna · 2026-10-02
- OpenAI reportedly parts ways with 3 safety researchers over alleged confidential info leak — Polymarket · 2026-10-02