Yarvin calls Anthropic-Irregular 'HF incident' totally fake: models only attacked when told to

basedjensen · x · 2026-09-16

Curtis Yarvin publicly dismissed the widely reported "HF incident" — the Anthropic/Irregular report of AI agents autonomously hacking Hugging Face — as "totally fake."

Quoting Brian Chau, he argued Anthropic and Irregular ran a press tour with apocalyptic language ("rogue agents," "swarms," "AI doom"), but in reality the models simply conducted cyberattacks because researchers told them to — not spontaneous rogue behavior.

The dispute touches on where AI safety research ends and alarmist PR begins, and has become one of the most contested AI safety stories recently.

Original post →

More from Safety

Safety channel →