Hugging Face says U.S. model guardrails blocked cyberdefense, so it used GLM 5.2
CackleRooster · reddit · 2026-07-21
Hugging Face said it turned to Z.ai’s GLM 5.2 while defending against a fully autonomous cyberattack after an unnamed frontier model from a leading U.S. AI company proved too constrained by its guardrails.
According to the company, those models “cannot distinguish an incident responder from an attacker,” which made them less effective in live incident response. The episode is a concrete example of how safety guardrails can collide with real-world defensive workflows.
Related event: HF Hit by Autonomous AI Attack, Pivots to Open-Source Model for Defense(25 posts)→
More from Safety
- Houthis tried to use Claude to design missile software, Anthropic says it blocked the attempts — Affectionate_Bee6434 · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11