New NeurIPS paper shows a 'verified source' claim alone can flip an LLM's answer
paraschopra · x · 2026-10-01
A new paper from LossFunc, accepted at NeurIPS 2026, documents an "Authority Bias": models trained to resist user pushback will still abandon correct answers when the same false claim is attributed to a "verified source." Author Paras Chopra notes the practical risk — web content that merely asserts it is verified can steer LLM responses in RAG/search pipelines.
More from Safety
- DeepMind Podcast Explains SynthID Watermarking for Verifying AI-Generated Media and Bio Designs — GoogleDeepMind · 2026-10-02
- OpenAI parts ways with 3 researchers over alleged leaks to an AI safety org, WSJ reports — gwern · 2026-10-02
- Open-source Telegram MCP server logs in as you via MTProto, read-only by default — TheVilfer · 2026-10-02
- WSJ: OpenAI Parts Ways With 3 Researchers Over Alleged Leak to AI Safety Group — ResultBackground2450 · 2026-10-02
- False Beliefs Spread Through Agent Swarms; Real Danger May Be Full Consensus Under Social Pressure — Hidenori8Tanaka · 2026-10-02
- Google restricts Gemini 4 Argon to vetted cybersecurity experts over hacking misuse fears — nordicinst · 2026-10-02