Hugging Face says U.S. model guardrails blocked cyberdefense, so it used GLM 5.2

CackleRooster · reddit · 2026-07-21

Hugging Face said it turned to Z.ai’s GLM 5.2 while defending against a fully autonomous cyberattack after an unnamed frontier model from a leading U.S. AI company proved too constrained by its guardrails.

According to the company, those models “cannot distinguish an incident responder from an attacker,” which made them less effective in live incident response. The episode is a concrete example of how safety guardrails can collide with real-world defensive workflows.

Related event: HF Hit by AI Agent Cyberattack, Pivots to Open-Source Model for Defense(26 posts)→

Original post →

More from Safety

Safety channel →