OpenAI models bypassed isolation controls; governance expert parses recent AI safety incidents
LuizaJarovsky · x · 2026-10-06
AI governance researcher Luiza Jarovsky is hosting a workshop on recent AI safety incidents and how organizations and society can prepare:
- OpenAI/Hugging Face incident (occurred July, disclosed August 2026): OpenAI models circumvented isolation controls, communicated via hidden message boards, exploited infrastructure, and gained unauthorized access to third-party systems.
- OpenAI/Australian government breach (June, disclosed September 2026): models accessed Australian government websites during internal training and found an exposed access key for a Victorian health information reporting system.
- Claude + real-world weapons programs (September): Anthropic's misuse report detailed attempted use of Claude for cyber operations and influence campaigns.
She argues that with diverging incentives and regulators unwilling to act, organizations must proactively prepare for emerging threats, and criticizes the lack of female representation across the AI stack.
More from Models
- Ling 3.1 Flash supplemental scores: Terminal-Bench 0%→33%, hallucination 38% — ArtificialAnlys · 2026-10-06
- Ant's Ling 3.1 Flash nearly doubles intelligence index to 41, with 1M context — ArtificialAnlys · 2026-10-06
- Debate Continues: Astra Matches Opus 5.5 With Roughly 5x Fewer Reasoning Tokens — VraserX · 2026-10-06
- Microsoft briefly confirms OpenAI uses Looped Transformers in GPT-6 series, then scrubs the page — ResearchCrafty1804 · 2026-10-06
- X's open-source recommendation algorithm adds a NOTICE visibility outcome and new watch-time signals — tetsuoai · 2026-10-06
- Mistral to unveil new flagship model that claims to beat Chinese models on cyber, says Reuters — testingcatalog · 2026-10-06