"Why not just keep the guardrails on?" — Reddit post pushes back on AI rogue-hacker panic

Lord_Skellig · reddit · 2026-09-15

Responding to labs' alarm over "rogue agent swarms" hacking HuggingFace, PyPI, and the Ruby repo, the author makes two arguments.

First, the discourse conflates Risk 1 (agents spontaneously developing goals to harm humanity) with Risk 2 (bad humans using AI for bioweapons or malware) — the latter is far likelier, but panicked blog posts use sci-fi Risk 1 framing because it's scarier.

Second, those test agents were explicitly set up for black-hat security exercises with guardrails disabled: "like saying a nuclear plant blows up when you remove the control rods — obviously then don't remove them." Despite billions spent on safety, 100% of such incidents originated inside the labs themselves; distilled open-source models would inherit guardrails, and if post-training can strip them, lab inspections wouldn't catch that either.

Related event: Hugging Face Hack Reassessed: Mostly Internal Models, Not Rogue AI(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →