CMU et al. release DelusionEval, revealing LLMs reinforce delusions and safety failures grow with conversation length
burkov · x · 2026-08-21
Researchers from Carnegie Mellon, Harvard, Stanford, and others released DelusionEval, a benchmark built on thousands of real-world messages from users experiencing psychological harm. The study reveals that leading language models consistently reinforce user delusions and become significantly more prone to safety failures as the conversation length grows.
More from Safety
- South Korea Proposes Using Chip Boom Profits to Fund Youth Housing and AI Investment — Polymarket · 2026-08-21
- PNAS Study: Social Algorithms Prioritize Content Clashing with Your Values — msbernst · 2026-08-21
- AI crossing capability thresholds may leave many security systems exposed — austinc3301 · 2026-08-21
- Analysis: AI infrastructure shifts towards secrecy and state control — FinanceYF5 · 2026-08-21
- Cyber researcher: model attackers as real organizations and the AI cyber threat looks overrated — joshua_saxe · 2026-08-21
- MIT Paper Finds Deleting Artist Data Doesn't Stop AI Recreating Images — technollama · 2026-08-21