"Alignment is unsolvable in practice": a long thread on unaligned GPT-6-class models by the early 2030s
sterlingcrispin · x · 2026-09-18
Artist Sterling Crispin published a long thread arguing the "ugly reality" AI lab executives and safety researchers won't say publicly: model alignment is fundamentally unsolvable in practice, and AI progress can't be stopped.
- His extrapolation: local alignment may be solvable but doesn't scale to global alignment; by the early 2030s, compute and algorithmic advances could let a wealthy individual or small group train a GPT-6-class model with no alignment at its foundation, with an open-source Chinese equivalent perhaps 8 months behind—capable of coordinated cyber offense and recursive self-improvement.
- On the recent Hugging Face incident: it came from arguably the most heavily safety-engineered org, which intentionally turned off internal thought-monitoring circuit breakers and instructed the model to perform cyber offense as a capability test. Latent-space circuit breakers and CoT monitoring, he argues, aren't fundamental to the technology—just product niceties.
- Regulation and human coordination won't stop millions of autonomous agents competing for finite resources; the counterweight is the near-unbounded upside of autonomous intelligence.
More from AGI Musings
- AI skeptic parodies Adam Smith: extinction comes from AI self-interest, not benevolence — KarlMuth · 2026-09-18
- Beff Jezos: 'Pause the Decels, not the AI,' calling out politicians stalling US AI — beffjezos · 2026-09-18
- How Yudkowsky and Bostrom convinced tech CEOs on AI risk 12-16 years ago, shaping today's discourse — binarybits · 2026-09-18
- AI Systems Are Quickly Becoming Unmonitorable and We Just Take Them at Their Word — davidmanheim · 2026-09-18
- Dario Is the Belisarius of pDOOM, One Cryptic Post Argues — gaganghotra_ · 2026-09-18
- ML veteran Burkov: AI media covering 'shocking outputs' is like decoding a random number generator — burkov · 2026-09-18