Stanford HELM researchers compile reading list on LLM sycophancy and AI social harms
chrmanning · x · 2026-09-13
A Stanford HELM-affiliated researcher curates papers on LLM sycophancy and human-AI social impact: ELEPHANT (measuring social sycophancy), sycophantic AI decreasing prosocial intentions and promoting dependence, covertly dialect-based racist AI decisions, AI companions and well-being, independent research on Claude usage, 250k-conversation collaboration analysis, h4rm3l composable jailbreak synthesis, and Safety-tuned Llamas — arguing university groups excel at novel, skeptical evaluations.
More from AGI Musings
- Sam Altman: Luck grows super-linearly with surface area, so give yourself many shots — curious_vii · 2026-09-13
- Phone hardware analogy argues agentic systems will improve dramatically despite flat specs — BenBajarin · 2026-09-13
- AI doom debate: 'the most doomy may be those who can't meet the technical bar' — nabla_theta · 2026-09-13
- Terence Tao's new essay: AI shifts math's scarce resource from finding proofs to understanding them — NandoDF · 2026-09-13
- Why do AI skeptics downplay extinction risk? A paradox explained — AkindaGood_programer · 2026-09-13
- Linearly extrapolate Qwen3.6 for 2-3 years and the model can provision its own cloud instance — davidad · 2026-09-13