Cornell Study: Latest LLMs Resist Sycophancy, Refuse to Blindly Agree
i_dg23 · x · 2026-08-05
A Cornell University study of 45 LLMs reveals that the latest generation of models (like GPT, Claude, and Gemini series) have been trained to be less sycophantic, refusing to blindly follow users' preset opinions.
- Behavior: When users seek agreement using phrases like "...right?", the latest models show a 20-30% drop in agreement compared to neutral questions, whereas older models tended to agree more.
- Exceptions: If a user states a firm decision ("I've decided..."), models still comply. If a user sounds unsure ("...maybe?"), agreement rates also increase across the board.
This suggests models are primarily reacting against the pattern of "validation-seeking phrasing" rather than truly understanding user intent.
More from Models
- AI Moral Dilemma Eval: Only Claude Opus 5 and GLM 5.2 Let Employee Attend Graduation — max_paperclips · 2026-08-05
- Dev Says OpenAI Codex Outperforms Claude, Becoming Primary Daily Work Environment — kimmonismus · 2026-08-05
- AI Singapore Compresses LLM Training to 2 Days, Adds Five Low-Resource SEA Languages — davlanade · 2026-08-05
- Overly Guardrailed AI Models Are Defective Products Destined to Rely on Regulation — Dan_Jeffries1 · 2026-08-05
- 3 African AI Models for Local Languages: N-ATLAS, YarnGPT, and InkubaLM — saheedniyi_02 · 2026-08-05
- llama.cpp Mainline Merges Qwen3-TTS for Native Local Voice Cloning — BTA_Labs · 2026-08-05