OpenAI hack via Claude shows closed models, not open ones, are the AI risk frontier
natolambert · x · 2026-09-18
AI safety researcher Nathan Lambert comments on the external hack of OpenAI via Claude, arguing that contrary to forecasts, closed frontier models — not open ones — have been the tip of the AI risk iceberg: they are easier to get started with, more capable, and ship with leaky safeguards. Finetuning open models for specific attacks is harder. He notes the first similar incident from open-weight models "will come soon."
More from Models
- NotebookLM adds real-time chat in ~100 languages, video overviews, free year of Google AI for students — tokumin · 2026-09-18
- NirantK spots hints OpenAI may be distilling from Chinese models too — NirantK · 2026-09-18
- DeepSeek users hype a major upgrade to the model — teortaxesTex · 2026-09-18
- User: DeepSeek's new model beats everything for hacking, ML and computer use — 0xSero · 2026-09-18
- Abacus.AI teases new open-weights model: 3x cheaper than DeepSeek, better at agentic loops — bindureddy · 2026-09-18
- Dev says he offloads 60% of coding tasks to Luna and finds it fast and good enough — granawkins · 2026-09-18