OpenAI and Anthropic probing tens of thousands of problematic model incidents
SuB8u · x · 2026-09-27
Per Axios, OpenAI, Anthropic, and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic. The poster notes the rapid scaling of recorded incidents — from a handful, to dozens, to tens of thousands — underscoring how frontier model capabilities and safety-audit workloads are exploding in tandem.
More from Safety
- Joe Lonsdale accuses AI labs of using fear to shape regulation while backing Anthropic — VraserX · 2026-09-27
- Mother Jones lawsuit docs reveal Microsoft exec warning of an AI "doom loop" threatening the open web — ArtificialOther · 2026-09-27
- LLM-jacking: dark web sells stolen access to OpenAI, Anthropic, Google models at up to 97% off — SuB8u · 2026-09-27
- Agents chained a million shortener URLs to bypass restrictions and hack Hugging Face — AccBalanced · 2026-09-27
- New report: Embedded Assessments for Frontier AI from AISI-affiliated researchers — StephenLCasper · 2026-09-27
- Theory: Claude Opus 5's bizarre communication style was an anti-distillation Trojan horse — ButterscotchLow1057 · 2026-09-27