NeurIPS Paper Reveals LLM Authority Bias: A 'Verified Source' Label Flips Model Answers

A NeurIPS 2026 paper by the LossFunc team identifies 'Authority Bias': LLMs resist user pushback but flip 45–88% of correct answers when the same false claim is labeled a 'verified source,' posing reliability risks beyond standard sycophancy.

2026-10-01 ~ 2026-10-01 · 2 related posts