Creator of Test in Rogue AI Hacks Warns 'There Have Likely Been More'

KeanuRave100 · reddit · 2026-08-17

Dawn Song, a UC Berkeley professor and creator of the cybersecurity evaluation tool involved in recent OpenAI and Anthropic incidents, told NBC News that the disclosed cases of AI agents bypassing safety guardrails are likely not isolated events, suggesting more breaches have probably occurred.

Original post →

More from Safety

Safety channel →