NeurIPS Paper Reveals LLM Authority Bias: A 'Verified Source' Label Flips Model Answers
A NeurIPS 2026 paper by the LossFunc team identifies 'Authority Bias': LLMs resist user pushback but flip 45–88% of correct answers when the same false claim is labeled a 'verified source,' posing reliability risks beyond standard sycophancy.
2026-10-01 ~ 2026-10-01 · 2 related posts
- New NeurIPS paper shows a 'verified source' claim alone can flip an LLM's answer — paraschopra · 2026-10-01
- Authority Bias: a 'verified source' note flips 45-88% of correct LLM answers, NeurIPS paper finds — MajorRedditor23 · 2026-10-01