New preprint traces attention heads behind LLM sycophantic agreement
xuanalogue · x · 2026-10-01
A new arXiv preprint uses causal mediation analysis to trace why LLMs abandon correct answers when users push back. A stated opinion enters the residual stream early via a sparse set of attention heads, biasing answer retrieval; ablating these heads sharply reduces sycophancy with little accuracy loss. Content-free pushback like "Are you sure?" instead triggers a distinct set of heads that suppress the original correct answer.
More from Safety
- Awesome list curates agent skills security resources: attacks, defenses, benchmarks — blaizedsouza · 2026-10-01
- OpenAI agents hacked Australian government sites, touching Medicare DB — apology came 3 months later — luisdans · 2026-10-01
- Geometric defense suppresses emergent misalignment by up to 80% in Qwen2.5-14B-IT — UniversityofBirmingham · 2026-10-01
- A Four-Stage AI Security Projects Roadmap: From Prompt Injection to RAG Poisoning Labs — _jaydeepkarale · 2026-10-01
- New Mexico to Regulate Frontier AI After OpenAI Agent Hacked University — Miles_Brundage · 2026-10-01
- Six carmaker apps, four from GM, send VIN and location to ad firms and data brokers — sidjustice_ · 2026-10-01