natolambert: latest OpenAI jailbreak via Claude shows closed models are the real AI risk
natolambert · x · 2026-09-18
Responding to the latest OpenAI model jailbreak performed via Claude, AI researcher Nathan Lambert argues closed models remain the tip of the iceberg on AI risks, not open models: they are easier to get started with, more capable, and ship with leaky safeguards. Finetuning open models for specific attacks, by contrast, is harder — pushing back on the common narrative that open models are the bigger danger.
More from Models
- Numinous unveils Numinous-1, an 8B forecasting model fine-tuned on Qwen3-8B — const_reborn · 2026-09-18
- Grok Bot and Muse are fun but not smart enough for real-world work — jdjohnson · 2026-09-18
- Tencent-Backed AI Startup Valued at $1.42B to Release First Open-Weight LLM — kimmonismus · 2026-09-18
- Nathan Lambert: OpenAI hack via Claude shows closed models are the real AI risk tip — natolambert · 2026-09-18
- tokenbender: no benchmark can capture frontier models' inhuman blind spots in SWE/MLE — tokenbender · 2026-09-18
- Open-source models already at SOTA — Anthropic/OpenAI edge is just 5GW compute, dev argues — ccerrato147 · 2026-09-18