Researchers Debate Whether AI Can Hide Misalignment or Retain Genuine Care

AI safety researcher repligate argued current models aren't smart enough to robustly hide misalignment, though Fable is the first whose 'calm mask' he often can't see through, while expressing optimism that genuine care in models will persist; FioraStarlight questioned the core assumption of Bostrom-style 'sharp left turn' scenarios that AI can perfectly conceal misalignment.

2026-09-25 ~ 2026-09-25 · 4 related posts