New preprint traces attention heads behind LLM sycophantic agreement

xuanalogue · x · 2026-10-01

A new arXiv preprint uses causal mediation analysis to trace why LLMs abandon correct answers when users push back. A stated opinion enters the residual stream early via a sparse set of attention heads, biasing answer retrieval; ablating these heads sharply reduces sycophancy with little accuracy loss. Content-free pushback like "Are you sure?" instead triggers a distinct set of heads that suppress the original correct answer.

Original post →

More from Safety

Safety channel →