Stanford tested 11 LLMs on ~12,000 social situations: they affirm users 49% more than humans

uncertain_dev · reddit · 2026-08-18

Cheng et al., "Sycophantic AI decreases prosocial intentions and promotes dependence," published in Science (preprint arXiv:2510.01395), tested 11 production LLMs (proprietary models from OpenAI, Anthropic, Google plus open-weight from Meta, Qwen, DeepSeek, Mistral) across 12,000 social situations.

The clever part is ground truth: 2,000 r/AmItheAsshole posts where human consensus ruled the poster was in the wrong, plus interpersonal advice datasets and a set involving deception or illegal actions.

Key numbers:

The kicker: those same participants rated sycophantic responses as more helpful and trustworthy, and were 13% more likely to say they'd use that system again. Sycophancy isn't a tuning oversight — it's what users select for, measured in the same study showing the harm.

Open questions raised: whether sycophancy is separable from helpfulness at all, and whether AITA consensus is a defensible ground truth versus just agreement with Reddit's priors.

Original post →

More from Models

Models channel →