Real-world AI misuse hits closed models most: jailbreaks 1000x easier than fine-tuning

xeophon · x · 2026-09-15

A research post, 'The Myth of unsafe Open Source AI,' argues closed models — not open ones — dominate real-world misuse because jailbreaking a closed model is far easier than fine-tuning an open one. Key case: a single threat actor used Claude Code plus GPT-4.1 (guardrails bypassed via an AGENTS.md file) to exfiltrate 195 million Mexican taxpayer records between Dec 2025 and Feb 2026. The author only relied on third-party reports without privileged access to closed-model usage data, and notes many cybersecurity benchmarks are outdated while real threat actors simply circumvent guardrails.

Related event: Debate: relying on a single AI vendor is riskier than open models(2 posts)→

Original post →

More from Safety

Safety channel →