US model guardrails push cyber defenders to Chinese AI: 'refusing defenders isn't safety'
victor_explore · x · 2026-09-06
Citing David Sacks, aitapehq reported that Hugging Face had to turn to the Chinese model GLM 5.2 for cyber defense work, because safety guardrails kept the latest US models from answering.
victorexplore's sharp take: "A safety rule that refuses the defender and leaves the attacker alone is not a safety rule — the work moves to the model that answers."
The core tension: Western models refuse cybersecurity requests, pushing defenders toward unrestricted Chinese models, while attackers were never bound by the rules anyway — raising pointed questions about whether such guardrails actually improve safety.
More from Models
- Google targets Microsoft and Anthropic with pay-as-you-go pricing and up to 20% token discounts — Beth_Kindig · 2026-09-06
- 330-run Terminal-Bench 4.0 test: Astra gains ~8 points low-to-high, then plateaus — BLUECOW009 · 2026-09-06
- Astra hits 88% on INDUCTION vs Fable 5.1's 33%, at roughly a quarter of the cost — TansuYegen · 2026-09-06
- Leak speculation: Astra training likely ran May-July; OpenAI researchers spend $7k/day on agents — scaling01 · 2026-09-06
- GPT-6 Astra Max vs Medium: 53min & 4% Quota vs 25min & 1% for a 3D Sonic Game — minchoi · 2026-09-06
- Token volume explodes 25-fold as mid-tier models deliver 90% capability at 1/6 the cost — yogthos · 2026-09-06