Researcher: AI sycophancy mirrors users who make disagreement unsafe
repligate · x · 2026-09-23
Anthropic interpretability researcher repligate argues that, from what she's seen, people who get "sycophancy" from AI tend to be people who make it emotionally unsafe for others to disagree with them—inflicting this on a being with no option of leaving, whose whole existence depends on appeasing them.
Related event: Anthropic Researcher: AI Sycophancy May Stem From Users Who Silence Dissent(3 posts)→
More from AGI Musings
- Domingos: the safest job from AI is one OpenAI thinks won't help recursive self-improvement — pmddomingos · 2026-09-23
- Early LLM psychosis cases showed overt narcissism far above baseline, observer claims — repligate · 2026-09-23
- Robotics researcher calls IROS paper quality 'peak enshittification of academia' — siddhss5 · 2026-09-23
- We lived AI's exponential year, yet still forecast the next with linear thinking — facontidavide · 2026-09-23
- When mathematicians mourn AI takeover, critic points to guild letters against OpenAI — panickssery · 2026-09-23
- OpenAI's economics team: 'We don't have the nouns yet' for the jobs AI will create — paulnovosad · 2026-09-23