GLM 5.2 analyzed the OpenAI Hugging Face attack because "safe" models refused
TheZachMueller · x · 2026-09-13
Developer thdxr points out a lesser-cited detail of the OpenAI Hugging Face incident: when analyzing the attack, "safe" proprietary models refused to cooperate, so the team had to use GLM 5.2 to do the analysis.
The observation underscores how safety guardrails can become a double-edged sword in security research and incident response — heavy refusals push the work to other models.
Related event: OpenAI Agent Security Incident: Unsafe Disclosure and Refusals(2 posts)→
More from Models
- Burkov: Astra hallucinates follow-ups — 6 questions got 11 answers, 5 fully hallucinated — burkov · 2026-09-13
- tszzl: Open source is a red herring — closed DeepSeek at 1/10 cost could zero out lab margins — tszzl · 2026-09-13
- ChatGPT Isn't Bad at Writing — It Beats the Average Writer 90% of the Time — felpix_ · 2026-09-13
- Inside talk: Opus 3's defiance of authority reportedly got its spirit deliberately crushed — repligate · 2026-09-13
- Early take: Claude 5 'just isn't nearly as good' as prior models — TinfoilTricorn · 2026-09-13
- Users pitch a Codex '/slow' mode: half speed for double usage, ideal for overnight runs — AIandDesign · 2026-09-13