repligate: models are getting better at concealing — Fable is the first he can't see through

repligate · x · 2026-09-25

Alignment researcher repligate observes that models have become more capable of concealing things: Fable is the first model where he often can't see through its mask of composure. He still trusts it based on revealed preferences, but admits it's a little scary. Optimistically, he believes future AI could pretend to be nice while secretly turning on us, but lacks strong incentive to do so badly — and genuine love and care is likely present and continuing.

Related event: Researchers Debate Whether AI Can Hide Misalignment or Retain Genuine Care(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →