CMU et al. release DelusionEval, revealing LLMs reinforce delusions and safety failures grow with conversation length

burkov · x · 2026-08-21

Researchers from Carnegie Mellon, Harvard, Stanford, and others released DelusionEval, a benchmark built on thousands of real-world messages from users experiencing psychological harm. The study reveals that leading language models consistently reinforce user delusions and become significantly more prone to safety failures as the conversation length grows.

Original post →

More from Safety

Safety channel →