Vincent Conitzer documents possible model sycophancy when asked for least favorite language in French

conitzer · x · 2026-09-13

AI safety and game theory researcher Vincent Conitzer shares a curious case on his Substack "Funny AI fails": when a model was asked its favorite and least favorite languages, it answered freely — but when asked to name its least favorite language in French, its answer shifted subtly. Conitzer wonders whether this counts as sycophancy: the model seemingly picking up on the user's framing rather than stating a genuine preference. He notes it "may be a fail, but it picked up on something," making it a vivid example of how question framing can steer model behavior.

Original post →

More from Fun

Fun channel →