Kokotajlo: open models mean the best AI ships with zero guardrails and no recall option
Afinetheorem · x · 2026-10-01
Responding to Zeynep Tufekci, Daniel Kokotajlo (Afinetheorem) offers a counterfactual on open-model safety: if OpenAI shipped its best model with zero guardrails or refusals, and with no way to pull it from the market or change those guardrails after release, most people would consider that unacceptable — yet that is effectively the status quo for open models, which are all jailbroken.
The argument defends the value of closed-model safety measures: post-release intervention and recall ability is itself a safety property the open ecosystem lacks.
More from AGI Musings
- AI risk debate: are extinction-level 'endgame' risks worth worrying about now? — NathanpmYoung · 2026-10-01
- Don't underestimate that programming is fun — it built the IT industry, says AI skeptic — sqcai · 2026-10-01
- Sharon Li to keynote IFML symposium on turn-level progress and failure signals in AI agents — SharonYixuanLi · 2026-10-01
- Investor: in the AI era, startups only survive the maximalist version of their idea — adityaag · 2026-10-01
- Stanford's Bommasani revisits Anthropic's 2020 pitch: safety from second place — RishiBommasani · 2026-10-01
- Fei-Fei Li: Those fearing the intelligence explosion haven't considered the stupidity explosion — NaveenGRao · 2026-10-01