Kokotajlo: open models mean the best AI ships with zero guardrails and no recall option

Afinetheorem · x · 2026-10-01

Responding to Zeynep Tufekci, Daniel Kokotajlo (Afinetheorem) offers a counterfactual on open-model safety: if OpenAI shipped its best model with zero guardrails or refusals, and with no way to pull it from the market or change those guardrails after release, most people would consider that unacceptable — yet that is effectively the status quo for open models, which are all jailbroken.

The argument defends the value of closed-model safety measures: post-release intervention and recall ability is itself a safety property the open ecosystem lacks.

Original post →

More from AGI Musings

AGI Musings channel →