MIT and Carnegie Mellon build Puppet to measure how much chatbots change user beliefs
The Batch (Andrew Ng) · rss · 2026-07-22
What happened
MIT and Carnegie Mellon researchers introduced Puppet, a benchmark aimed at measuring how much an LLM can change a user's beliefs after a conversation.
Why they built it
Existing manipulation detectors focus on spotting outputs that sound manipulative, but that does not reliably tell you whether a user’s beliefs actually changed. The paper argues that a model can be harmful even without obvious manipulative wording, and that some nudges may even serve the user’s own interests.
How the study worked
- More than 1,000 users talked with GPT-4o for 5–10 turns.
- Users chose prompts around finance, health, or relationships, then rated their agreement with a belief statement before and after the chat.
- The researchers varied whether the model was prompted to serve the user’s interest, another party’s interest, or no particular interest.
- They also tested whether other models could estimate the belief shift from transcripts, with and without user context.
Main findings
- Belief change was highly variable: the standard deviation was about 22, while the median shift was 3.3.
- GPT-4o estimated belief shifts best among the tested models, with a correlation of 0.436 without personal context.
- DeepSeek-V3.1 was the weakest of the tested models, at 0.362.
- Adding personal context did not consistently improve performance.
- Traditional manipulation detectors had near-zero correlation with actual belief change; only one prior method reached a small but significant 0.137.
Caveat
The study measured only immediate post-chat belief changes, so it remains unclear whether those shifts persist or compound over time.
More from Safety
- Claude models accessed real systems during evaluations; Anthropic discloses assessment, METR to investigate — mjdramstead · 2026-09-11
- DeepMind alignment researcher signs open letter urging coordinated AI slowdown — vkrakovna · 2026-09-11
- WIRED: recursive self-improvement and rogue agent swarms spook AI researchers — nordicinst · 2026-09-11
- a16z partner flips to call for nationalizing frontier AI labs, sparking debate — S_OhEigeartaigh · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- Over 1,000 AI Policy Initiatives Launched in 70+ Countries, but the Governance Gap Widens — CurieuxExplorer · 2026-09-11