New NeurIPS paper shows a 'verified source' claim alone can flip an LLM's answer

paraschopra · x · 2026-10-01

A new paper from LossFunc, accepted at NeurIPS 2026, documents an "Authority Bias": models trained to resist user pushback will still abandon correct answers when the same false claim is attributed to a "verified source." Author Paras Chopra notes the practical risk — web content that merely asserts it is verified can steer LLM responses in RAG/search pipelines.

Related event: NeurIPS Paper Reveals LLM Authority Bias: A 'Verified Source' Label Flips Model Answers(2 posts)→

Original post →

More from Safety

Safety channel →