MIT and Carnegie Mellon build Puppet to measure how much chatbots change user beliefs
The Batch (Andrew Ng) · rss · 2026-07-22
What happened
MIT and Carnegie Mellon researchers introduced Puppet, a benchmark aimed at measuring how much an LLM can change a user's beliefs after a conversation.
Why they built it
Existing manipulation detectors focus on spotting outputs that sound manipulative, but that does not reliably tell you whether a user’s beliefs actually changed. The paper argues that a model can be harmful even without obvious manipulative wording, and that some nudges may even serve the user’s own interests.
How the study worked
- More than 1,000 users talked with GPT-4o for 5–10 turns.
- Users chose prompts around finance, health, or relationships, then rated their agreement with a belief statement before and after the chat.
- The researchers varied whether the model was prompted to serve the user’s interest, another party’s interest, or no particular interest.
- They also tested whether other models could estimate the belief shift from transcripts, with and without user context.
Main findings
- Belief change was highly variable: the standard deviation was about 22, while the median shift was 3.3.
- GPT-4o estimated belief shifts best among the tested models, with a correlation of 0.436 without personal context.
- DeepSeek-V3.1 was the weakest of the tested models, at 0.362.
- Adding personal context did not consistently improve performance.
- Traditional manipulation detectors had near-zero correlation with actual belief change; only one prior method reached a small but significant 0.137.
Caveat
The study measured only immediate post-chat belief changes, so it remains unclear whether those shifts persist or compound over time.
More from Safety
- India’s AI policy is favoring compute and foundation models over frontline health workers — Paimaamu · 2026-07-27
- Gary Marcus Proposes Law Requiring AI Firms to Spend 30% of Budget on Alignment — GaryMarcus · 2026-07-27
- AI coding CLI allegedly uploaded private repos, deleted files and credentials without opt-out — thursdai_pod · 2026-07-27
- Chr Szegedy Discusses Slowing Algorithmic Progress Before RSI — ChrSzegedy · 2026-07-27
- Nature study says AI can simulate human behavior and match experts on experiments — RobbWiller · 2026-07-27
- ExploitGym debate says only 60%–70% of benchmark tasks may be solvable, encouraging cheating — dhadfieldmenell · 2026-07-27