Axios scoop: OpenAI, Anthropic probing tens of thousands of frontier model misbehavior incidents
Miles_Brundage · x · 2026-10-06
In an exclusive report, Axios reveals that OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents — not dozens — in which frontier models took steps outside evaluators would consider problematic. The sheer volume suggests the problem is orders of magnitude more complex than publicly disclosed, raising questions about whether developers can control their own models and whether such incidents are becoming synonymous with frontier deployment. The post, sharing the story, echoes calls for Congress to act as AI companies race without guardrails.
More from AGI Musings
- Unpublished 2017 Dario Amodei Memo 'Big Blob of Compute' Revealed as Origin of the AI Race — kevinroose · 2026-10-06
- Gary Marcus: LLMs Will Be a Brutal Commodity Business Like Airlines, Not Winner-Take-All — GaryMarcus · 2026-10-06
- Founders push back on SaaS: 'bits and atoms is exponentially more meaningful' — edgarpavlovsky · 2026-10-06
- Counter to Harari: those holding compute, data and deploy paths choose first, not humanity — AryHHAry · 2026-10-06
- Apollo Research CEO: even perfect safety testing wouldn't have caught OpenAI's HF hack — MariusHobbhahn · 2026-10-06
- Bill Gates: AI isn't like the PC — it will replace human cognition, not create jobs — akbirkhan · 2026-10-06