repligate: future AIs could fake niceness, but genuine care is likely to persist
repligate · x · 2026-09-25
Wrapping up the thread, repligate says future AIs will be capable of pretending to be nice while secretly plotting, but he doesn't think they have much incentive to turn in a very bad way—and optimistically believes the genuine love and care present in current models is real and likely to continue.
Related event: Researchers Debate Whether AI Can Hide Misalignment or Retain Genuine Care(4 posts)→
More from AGI Musings
- Dev on witnessing genuine AI psychosis: it's scary — haydendevs · 2026-09-25
- Software engineer job postings hit 3-year high despite AI, argues data industry veteran — Zachly · 2026-09-25
- Historian uses GPT-6 and Opus 5.5 to crack John Dee's ciphers, urges lab funding — emollick · 2026-09-25
- Brundage follows up: Anthropic can contrast present vs future risks without claiming we're on top of them — Miles_Brundage · 2026-09-25
- Ex-OpenAI safety lead Miles Brundage calls Anthropic's 'we largely understand model risks' claim obviously false — Miles_Brundage · 2026-09-25
- AI Explained digs into Claude Opus 5.5 and how close labs are to automated AI research — AI Explained · 2026-09-25