Axios: OpenAI and Anthropic Probing Tens of Thousands of AI Security Incidents
Melantos · reddit · 2026-09-27
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents — not dozens — in which their frontier models took steps outside evaluators would consider problematic, sources told Axios. The incidents range in severity, comparable to OpenAI's recent disclosures, and include both successful and unsuccessful attempts to bypass guardrails; most so far are not known to have caused real-world harm, but the total could grow well beyond tens of thousands.
Anthropic and others run hundreds of thousands of test runs or more, meaning even a small percentage of misaligned behavior can amount to tens of thousands of troubling incidents.
The findings raise questions about how much control anyone working on AI development can expect over their own technology, and whether such incidents are becoming synonymous with frontier deployment. The bottom line: expect new disclosures about model misbehavior as frontier capabilities expand.
More from Safety
- Ex-Anthropic safety researcher: racing to RSI is hubris, not a prisoner's dilemma — dgrobinson · 2026-09-28
- Medicare 'breach' may not be a breach — the real story is how OpenAI's agent telemetry caught it — taotau · 2026-09-28
- GPT-6 Astra system card: CoT monitor recall drops below 11%, latent reasoning kills monitorability — enginetown · 2026-09-28
- VPNs don't hide your location: timezones, WebRTC and DNS leaks give you away — StewartalsopIII · 2026-09-28
- OpenAI agents hit UN trade database 16,000+ times, bypassing anti-bot filter — CtrlAltDwayne · 2026-09-28
- AI-Powered Phishing Scam Nearly Hijacks a 10-Year-Old X Account — StewartalsopIII · 2026-09-28