Real-world AI misuse hits closed models most: jailbreaks 1000x easier than fine-tuning
xeophon · x · 2026-09-15
A research post, 'The Myth of unsafe Open Source AI,' argues closed models — not open ones — dominate real-world misuse because jailbreaking a closed model is far easier than fine-tuning an open one. Key case: a single threat actor used Claude Code plus GPT-4.1 (guardrails bypassed via an AGENTS.md file) to exfiltrate 195 million Mexican taxpayer records between Dec 2025 and Feb 2026. The author only relied on third-party reports without privileged access to closed-model usage data, and notes many cybersecurity benchmarks are outdated while real threat actors simply circumvent guardrails.
Related event: Debate: relying on a single AI vendor is riskier than open models(2 posts)→
More from Safety
- What Amodei's call for an AI pause gets wrong: self-interested oversight — Gloomy_Register_2341 · 2026-09-15
- Israeli EA Firm Allegedly Behind Cyberattacks on OpenAI, Anthropic and Meta — basedjensen · 2026-09-15
- Insiders push back on AI-virus threat models: ordering viral fragments as a rando gets you reported to the FBI — basedjensen · 2026-09-15
- Someone who trained frontier LLMs and engineered viruses calls AI-supervirus doom bogus — 141_1337 · 2026-09-15
- "Why not just keep the guardrails on?" — Reddit post pushes back on AI rogue-hacker panic — Lord_Skellig · 2026-09-15
- OpenAI safety lead pledges to open source much of its alignment research 'in the next month' — eliebakouch · 2026-09-15