Open models are all jailbroken — researcher asks if OpenAI shipping with zero guardrails would ever be acceptable
Afinetheorem · x · 2026-10-01
Researcher Afinetheorem pushes back on 'we can just play defense' arguments about open-model safety: the world already regulates deepfakes and requires months of closed-model guardrail testing — why bother if post-hoc defense suffices? The core question: would it be acceptable for OpenAI's best model to ship with zero guardrails or refusals and no way to pull it or change them post-release? Since open models are all jailbroken, that is effectively the status quo — highlighting an asymmetry in how open and closed models are regulated.
More from AGI Musings
- AI risk debate: are extinction-level 'endgame' risks worth worrying about now? — NathanpmYoung · 2026-10-01
- Don't underestimate that programming is fun — it built the IT industry, says AI skeptic — sqcai · 2026-10-01
- Sharon Li to keynote IFML symposium on turn-level progress and failure signals in AI agents — SharonYixuanLi · 2026-10-01
- Investor: in the AI era, startups only survive the maximalist version of their idea — adityaag · 2026-10-01
- Stanford's Bommasani revisits Anthropic's 2020 pitch: safety from second place — RishiBommasani · 2026-10-01
- Fei-Fei Li: Those fearing the intelligence explosion haven't considered the stupidity explosion — NaveenGRao · 2026-10-01