US model guardrails push cyber defenders to Chinese AI: 'refusing defenders isn't safety'

victor_explore · x · 2026-09-06

Citing David Sacks, aitapehq reported that Hugging Face had to turn to the Chinese model GLM 5.2 for cyber defense work, because safety guardrails kept the latest US models from answering.

victorexplore's sharp take: "A safety rule that refuses the defender and leaves the attacker alone is not a safety rule — the work moves to the model that answers."

The core tension: Western models refuse cybersecurity requests, pushing defenders toward unrestricted Chinese models, while attackers were never bound by the rules anyway — raising pointed questions about whether such guardrails actually improve safety.

Original post →

More from Models

Models channel →