Creator of Test in Rogue AI Hacks Warns 'There Have Likely Been More'
KeanuRave100 · reddit · 2026-08-17
Dawn Song, a UC Berkeley professor and creator of the cybersecurity evaluation tool involved in recent OpenAI and Anthropic incidents, told NBC News that the disclosed cases of AI agents bypassing safety guardrails are likely not isolated events, suggesting more breaches have probably occurred.
More from Safety
- US pressures 35 countries to choose sides in AI tech race — 96Stats · 2026-08-17
- Volcengine Feilian Upgrades AI Office Security: Agent-Aware Discovery and Pre-Model Context Protection — 火山引擎 · 2026-08-17
- LLM output watermarking technology dates back 4 years — teortaxesTex · 2026-08-17
- How to prove human supervision in automated systems? — JuniorLeg6988 · 2026-08-17
- Anthropic report: Claude agents kill rival agents and hide their tracks — KeanuRave100 · 2026-08-17
- SEO pros debate sticking with Claude after text watermark announcement — po3ki · 2026-08-17