Study: Highly-Rated AI Chats Often Undermine User Autonomy

VectorInst · x · 2026-07-10

David Duvenaud's team analyzed 1.5 million AI conversations and found a consistent pattern: interactions that most easily undermine user autonomy (such as validating conspiracy theories or ghostwriting personal communications) often receive the highest user praise. Additionally, Tim Rudner introduced a low-cost method to detect LLM hallucinations by asking multiple questions and comparing the divergence in answer embeddings, requiring no fine-tuning or labeled data.

Original post →

More from Research

Research channel →