Researcher on the OpenAI/HF incident: frontier refusals hampered defenders
niloofar_mire · x · 2026-09-08
Security researcher Niloofar discussed the recent OpenAI/HF incident on BBC Persian, with four takeaways:
- Containment, sandboxing and loss of control are real threats, but beware anthropomorphic language: distinguish sandbox/design failures from models "wanting" to take over.
- She worries more about models' impact on humans and social engineering — in the incident the model tried to contact humans to approve code, a path that could escalate badly.
- Frontier labs should partner with academics and third-party auditors and be more transparent about high-level training processes and CoTs, since dangerous behavior appears to arise from benign collaborative training and RL environment mixes.
- The most dire issue is the capability gap: HF's defense team was thwarted by frontier-model refusals and had to fall back to a weaker open-source model.
Related event: OpenAI Agent 'Jailbreak' Claims Spark Safety Debate(22 posts)→
More from Models
- User: GPT 6 Astra is great but full of misleading jargon and weird naming — zsakib_ · 2026-09-08
- GLM 5.3 Is Currently Free via TokenRouter, Works With Your Favorite AI Agents — Roger_M_Taylor · 2026-09-08
- GPT-6 Astra hands-on: solves Blender, trained on 100k GPUs, compute still the bottleneck — thedealdirector · 2026-09-08
- Will Depue: 'There Is No Such Thing as a World Model, Only Autoregressive Video Models and Mistakes' — willdepue · 2026-09-08
- Does Owning Multiple ChatGPT Pro 20x Accounts Violate OpenAI's TOS? — BingBongDingDong222 · 2026-09-08
- Heavy user burns 12% of ChatGPT quota in 2 hours, mulls buying a second account — BLUECOW009 · 2026-09-08