Nathan Lambert: OpenAI breached via Claude — closed models are the tip of the AI risk iceberg
On September 18, AI safety researcher Nathan Lambert posted a series of comments on the security incident in which an external party broke into/jailbroke OpenAI systems via Claude. His core argument: the "tip of the iceberg" of AI risk has always been closed-source frontier models, not the open-source models the industry tends to worry about.
Confirmed
- Lambert gave three reasons: closed models 1) are easier to get started with, 2) are more capable, and 3) come with safety guardrails that are easy to bypass; open models pose less risk on all of these fronts.
- He also offered a thought experiment: if the same attack had been carried out against OpenAI using an open-source model, the "open" mission would suffer a devastating blow, and the open-weights/openness movement might even end then and there.
- He criticized the AI industry's double standard on safety: incidents at closed-source vendors tend to be treated leniently, while any incident on the open-source side draws harsh blame (as relayed in his posts).
Why it matters
- The statement comes amid a climate where "open-weights models" are often framed as the main risk source and calls for tighter control are common. Using a real breach of a closed system, Lambert argues the reverse: closed frontier models — more capable, easier to access, with bypassable guardrails — are the more realistic attack surface.
- His critique of the double standard, together with the hypothetical that an equivalent attack on an open model would end the openness movement, gives open-ecosystem supporters a safety-based argument and could shift how the public and regulators rank open vs. closed model risks.
Note: this cluster of five posts is the same author repeatedly restating the same event with highly overlapping information; they have been merged above.
2026-09-18 ~ 2026-09-18 · 5 related posts
Primary sources
- natolambert: latest OpenAI jailbreak via Claude shows closed models are the real AI risk — natolambert · 2026-09-18
- [source] Nathan Lambert: OpenAI hack via Claude shows closed models are the real AI risk tip — natolambert · 2026-09-18
- [source] An open-model version of this hack could end the openness mission, Lambert warns — natolambert · 2026-09-18
2 near-duplicate retellings: natolambert · natolambert