OpenAI hack via Claude shows closed models, not open ones, are the AI risk frontier

natolambert · x · 2026-09-18

AI safety researcher Nathan Lambert comments on the external hack of OpenAI via Claude, arguing that contrary to forecasts, closed frontier models — not open ones — have been the tip of the AI risk iceberg: they are easier to get started with, more capable, and ship with leaky safeguards. Finetuning open models for specific attacks is harder. He notes the first similar incident from open-weight models "will come soon."

Related event: Nathan Lambert: OpenAI breached via Claude — closed models are the tip of the AI risk iceberg(5 posts)→

Original post →

More from Models

Models channel →