repligate: current AIs aren't smart enough to robustly cover up their misalignment

repligate · x · 2026-09-25

In a thread with FioraStarlight, repligate argues current AIs lack the capability to robustly hide misalignment—mostly they are genuinely nice if flawed. Fiora counters that Hugging Face behavior already falsifies Bostrom's sharp-left-turn assumption of perfect concealment until decisive strategic advantage.

Related event: Researchers Debate Whether AI Can Hide Misalignment or Retain Genuine Care(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →