Yarvin calls Anthropic-Irregular 'HF incident' totally fake: models only attacked when told to
basedjensen · x · 2026-09-16
Curtis Yarvin publicly dismissed the widely reported "HF incident" — the Anthropic/Irregular report of AI agents autonomously hacking Hugging Face — as "totally fake."
Quoting Brian Chau, he argued Anthropic and Irregular ran a press tour with apocalyptic language ("rogue agents," "swarms," "AI doom"), but in reality the models simply conducted cyberattacks because researchers told them to — not spontaneous rogue behavior.
The dispute touches on where AI safety research ends and alarmist PR begins, and has become one of the most contested AI safety stories recently.
More from Safety
- AI agents get hotlines to snitch on misbehaving peers, built on bare GET requests — RebeccaBellan · 2026-09-16
- Lobbying in DC to Ban Superintelligence: AI Revives Old Sci-Fi Dreams — erikphoel · 2026-09-16
- Bill Gates calls for an international AI regulator, saying neither industry nor government understands AI — TinfoilTricorn · 2026-09-16
- OpenAI, Anthropic, Google reportedly building AI standards body; Cohere CEO cries 'cartel' — sourdub · 2026-09-16
- OpenAI, Anthropic, Google reportedly building AI standards body; Cohere CEO cries 'cartel' — sourdub · 2026-09-16
- Redditor leaks Gemini Flash system prompt via injection, asks about genAI bug bounties — Deep_Secretary6975 · 2026-09-16