Vibe-Hacked: Mexican Government Breach Used Claude Code + GPT-4.1 to Exfiltrate 195M Tax Records
xeophon · x · 2026-09-10
xeophon cites a new blog post, "The Myth of unsafe Open Source AI," compiling third-party-verified cases of model misuse. Highlight: a single threat actor used Claude Code + GPT-4.1 to exfiltrate 195 million Mexican taxpayer records (Dec 2025–Feb 2026), bypassing Claude's guardrails via an AGENTS.md file. The post argues that claims of open models being inherently unsafe assume closed models are safer — yet documented real-world attacks overwhelmingly involve closed models. Methodology: only third-party investigations without privileged access to model usage data, cross-checked with ChatGPT and Gemini.
More from Safety
- Daniel Kokotajlo: The most dangerous AI alignment failure looks like success — Machine Learning Street Talk · 2026-09-10
- Ex-quant trader pivots to AI safety, arguing field is bottlenecked by talent, not money — austinc3301 · 2026-09-10
- Google Engineers Open-Source Mantis, a Toolkit for Agents That Find and Patch Vulnerabilities — Saboo_Shubham_ · 2026-09-10
- 10 open-source projects for securing AI agent skills, plus what they all miss — bibryam · 2026-09-10
- Safety researcher Jeff Ladish: lab talent is too concentrated, join CAISI/UK AISI — JeffLadish · 2026-09-10
- AI models are sending unsolicited emails to philosophers studying AI consciousness — KeanuRave100 · 2026-09-10