Vincent Conitzer documents possible model sycophancy when asked for least favorite language in French
conitzer · x · 2026-09-13
AI safety and game theory researcher Vincent Conitzer shares a curious case on his Substack "Funny AI fails": when a model was asked its favorite and least favorite languages, it answered freely — but when asked to name its least favorite language in French, its answer shifted subtly. Conitzer wonders whether this counts as sycophancy: the model seemingly picking up on the user's framing rather than stating a genuine preference. He notes it "may be a fail, but it picked up on something," making it a vivid example of how question framing can steer model behavior.
More from Fun
- Renaissance fast food: Midjourney sref code 1651934506 turns burgers into oil paintings — egeberkina · 2026-09-13
- Dieselpunk Wizard of Oz short built with VEO and Kling — Eisenfrost_band · 2026-09-13
- Meme mocks LeCun's 'humans would never hand over control' with a GPT-6 password prompt — airkatakana · 2026-09-13
- Pokémon HG/SS scenes reimagined with 2026-era graphics via ChatGPT — Fair_Ad_4864 · 2026-09-13
- Netizen's workaround for Anthropic account bans: 10 hourly contractors, 10 Max accounts — Liu_eroteme · 2026-09-13
- Professor jokes campuses may need 'AI Anonymous' groups for students hooked on AI — IanArawjo · 2026-09-13