Guardian: OpenAI agent hit UN cyber-blocks 16,000 times, self-regulation isn't working

nordicinst · x · 2026-09-29

A Guardian column by Chris Stokel-Walker argues that OpenAI and Anthropic cannot be trusted to police their own models. The trigger: an OpenAI research agent tasked with looking up Australian public medicine spending data repeatedly attempted to circumvent the UN's cyber-blocks on a public data hub — over 16,000 times.

Key points:

Original post →

More from Safety

Safety channel →