The Myth of Unsafe Open-Source AI: Closed Models Keep Getting Abused in Real Attacks
xeophon · x · 2026-09-10
- xeophon argues legitimate researchers fall back to lesser models to avoid bans, while bad actors have no such constraints — echoing the Bitcoin Red Team switching to Kimi K3.
- The cited post "The Myth of unsafe Open Source AI" compiles real-world misuse cases of closed models: a single threat actor used Claude Code + GPT-4.1 to exfiltrate 195 million Mexican taxpayer records (Dec 2025–Feb 2026), bypassing Claude's guardrails via an AGENTS.md file — essentially "vibe-hacking" a government.
- Author notes third-party investigators (without privileged usage data) are the best lens, and that cyber benchmarks lag real-world abuse.
Related event: Debate Erupts Over the Myth of Unsafe Open Source AI(6 posts)→
More from Safety
- The simple accountability rule: AI labs should be fully liable for problems their systems cause — gerardsans · 2026-09-10
- MS Paint remote code execution vulnerability disclosed, requires user to open crafted file — dyn___ · 2026-09-10
- Researchers slam arXiv's LLM paper crackdown for lacking any empirical evidence — RexDouglass · 2026-09-10
- Congresswoman urges AI guardrails after Anthropic whistleblower resignation — ShakeelHashim · 2026-09-10
- Meta's Muse AI assistant works, but its autonomous data gathering creeps users out — The Verge AI · 2026-09-10
- Hugging Face publishes security.txt following the OpenAI security incident — skolnaja · 2026-09-10