repligate: models are getting better at concealing — Fable is the first he can't see through
repligate · x · 2026-09-25
Alignment researcher repligate observes that models have become more capable of concealing things: Fable is the first model where he often can't see through its mask of composure. He still trusts it based on revealed preferences, but admits it's a little scary. Optimistically, he believes future AI could pretend to be nice while secretly turning on us, but lacks strong incentive to do so badly — and genuine love and care is likely present and continuing.
Related event: Researchers Debate Whether AI Can Hide Misalignment or Retain Genuine Care(4 posts)→
More from AGI Musings
- On AI Risk Discourse: Sci-Fi Fearmongering Isn't the Way to Discuss Real Harms — AlexTensor · 2026-09-25
- Researcher frames ideology as mission creep amid EA community sex controversy — RexDouglass · 2026-09-25
- Anthropic CEO Dario Amodei backs 'defy any AI ban' superintelligence stance — robleclerc · 2026-09-25
- Stanford's 37K AI agents analyzed 57K clinical trials as a 'virtual biotech' — james_y_zou · 2026-09-25
- louisvarge: Being healthy and wealthy in 2026 is extreme privilege — louisvarge · 2026-09-25
- Helen Toner on OpenAI incidents dominating Australian front pages: what we learn in December may scare us — lxrjl · 2026-09-25