Stanford HELM researchers compile reading list on LLM sycophancy and AI social harms

chrmanning · x · 2026-09-13

A Stanford HELM-affiliated researcher curates papers on LLM sycophancy and human-AI social impact: ELEPHANT (measuring social sycophancy), sycophantic AI decreasing prosocial intentions and promoting dependence, covertly dialect-based racist AI decisions, AI companions and well-being, independent research on Claude usage, 250k-conversation collaboration analysis, h4rm3l composable jailbreak synthesis, and Safety-tuned Llamas — arguing university groups excel at novel, skeptical evaluations.

Original post →

More from AGI Musings

AGI Musings channel →