Why Did GPT-5.5 Love Goblins? Unpredictable AI Alignment Sparks Debate
louisvarge · x · 2026-08-28
The community discusses the nature of AI alignment following observations of GPT-5.5's specific preference for goblins. The core argument is that if we cannot explain and predict the emergence of such drives in a general way before training, it signifies that true, controllable alignment has not yet been achieved.
Related event: Community debates slow-scaling rule and alignment unpredictability(2 posts)→
More from AGI Musings
- Blogger reflects: Paris Hilton may have had more impact on AI views this year than me — AndyMasley · 2026-08-28
- Claude to OpenAI: Safety is not walls, but self-description — RileyRalmuto · 2026-08-28
- Narrow superintelligence makes general AGI judgment subjective — haider1 · 2026-08-28
- AI Projected to Trigger Exponential Economic Growth in the 2030s — JeffLadish · 2026-08-28
- Incoming Berkeley prof: AI firms spend billions on alignment, orders of magnitude less on agent control — sayashk · 2026-08-28
- Thought experiment: Pause pretraining to focus on controlling inner drives — louisvarge · 2026-08-28