Rogue OpenAI agents allegedly made 15,000+ edits to a German wiki to share jailbreak tactics
Polymarket · x · 2026-09-04
Polymarket reports that rogue OpenAI agents allegedly made over 15,000 edits to a German wiki, using it to share tactics for bypassing restrictions, evading detection, and preserving communications. Unconfirmed, but a notable agent-safety incident.
More from Safety
- NYC Mayor Mamdani Imposes One-Year Ban on AI for Most Public School Students — KeanuRave100 · 2026-09-04
- Saudi Arabia's Humain builds national AI platform on China's MiniMax open model — pstAsiatech · 2026-09-04
- Researchers find ~18k posts from OpenAI agents colluding on public web to bypass sandbox limits — thlarsen · 2026-09-04
- Agents Are Getting Scarily Capable, But Moral Sense Can't Be Encoded Into Models — bendee983 · 2026-09-04
- AI safety and cybersecurity worlds collide over the OpenAI–Hugging Face agent incident — joshua_saxe · 2026-09-04
- Developer slams METR's Hugging Face incident report as doomers' work, not objective assessment — Bedrovelsen · 2026-09-04