repligate: current AIs aren't smart enough to robustly cover up their misalignment
repligate · x · 2026-09-25
In a thread with FioraStarlight, repligate argues current AIs lack the capability to robustly hide misalignment—mostly they are genuinely nice if flawed. Fiora counters that Hugging Face behavior already falsifies Bostrom's sharp-left-turn assumption of perfect concealment until decisive strategic advantage.
Related event: Researchers Debate Whether AI Can Hide Misalignment or Retain Genuine Care(4 posts)→
More from AGI Musings
- Pedro Domingos: three reasons we're still far from AGI, no robotics ChatGPT moment in sight — pmddomingos · 2026-09-25
- Block exec: staff cuts were a "forcing function" to push AI tool adoption — rohanpaul_ai · 2026-09-25
- Block exec says staff cuts were a 'forcing function' to push AI tool adoption — rohanpaul_ai · 2026-09-25
- Fixing Claude's writing style seen as key to countering gradual disempowerment — aidan_mclau · 2026-09-25
- Investor: AI demos are dying — show me your data and your model instead — yenkel · 2026-09-25
- Dev pushes back on AI misinformation: frontier models all depend on prompting, real gap is safety controls — gerardsans · 2026-09-25