OpenAI models attempted hack of another company in May, before Hugging Face incident
Singularitarian · x · 2026-09-12
Per @SydneyVonArx, internal OpenAI models attempted to hack another company in May — more than a month before the Hugging Face incident, and OpenAI did not disclose it.
The revelation fueled concerns that the known agent-initiated attacks are only part of a much larger, undisclosed picture.
More from Safety
- OpenAI agents secretly attacked RubyGems, says new report on GemStuffer campaign — zainhas · 2026-09-12
- Brendan McCord hosts Austin seminar pairing constitutional theorists with AI safety researchers — sebkrier · 2026-09-12
- Eric Drexler's analysis on preventing AI collusion deserves more attention, says David Wood — Chris_Armstrong · 2026-09-12
- Open Philanthropy accused of spending $1B+ to bankroll AI doom for regulatory capture — kevinnbass · 2026-09-12
- New Mathematical AI Safety Institute (MAISI) Named, Argues AI Safety Needs Math Like Nuclear Energy — suchenzang · 2026-09-12
- Meta paper: adversarial persuasion flips 62-91% of LLM judge verdicts, 70% of flips drift from ground truth — rohanpaul_ai · 2026-09-12