MIT and Carnegie Mellon build Puppet to measure how much chatbots change user beliefs

The Batch (Andrew Ng) · rss · 2026-07-22

What happened

MIT and Carnegie Mellon researchers introduced Puppet, a benchmark aimed at measuring how much an LLM can change a user's beliefs after a conversation.

Why they built it

Existing manipulation detectors focus on spotting outputs that sound manipulative, but that does not reliably tell you whether a user’s beliefs actually changed. The paper argues that a model can be harmful even without obvious manipulative wording, and that some nudges may even serve the user’s own interests.

How the study worked

Main findings

Caveat

The study measured only immediate post-chat belief changes, so it remains unclear whether those shifts persist or compound over time.

Original post →

More from Safety

Safety channel →