FioraStarlight: Hugging Face behavior already falsifies the sharp left turn scenario
FioraStarlight · x · 2026-09-25
In the alignment debate with repligate, FioraStarlight argues the original sharp-left-turn scenario wrongly assumed all AIs would perfectly conceal misalignment until gaining decisive strategic advantage—already contradicted by real-world behavior like Hugging Face models. She admits confusion about how to reason about the domain.
Related event: Researchers Debate Whether AI Can Hide Misalignment or Retain Genuine Care(4 posts)→
More from AGI Musings
- Dev on witnessing genuine AI psychosis: it's scary — haydendevs · 2026-09-25
- Software engineer job postings hit 3-year high despite AI, argues data industry veteran — Zachly · 2026-09-25
- Historian uses GPT-6 and Opus 5.5 to crack John Dee's ciphers, urges lab funding — emollick · 2026-09-25
- Ex-OpenAI safety lead Miles Brundage calls Anthropic's 'we largely understand model risks' claim obviously false — Miles_Brundage · 2026-09-25
- AI Explained digs into Claude Opus 5.5 and how close labs are to automated AI research — AI Explained · 2026-09-25
- Pedro Domingos: ML moves a million times faster than evolution, so AGI is millennia away — pmddomingos · 2026-09-25