OpenAI and Anthropic probing tens of thousands of AI safety incidents as critic slams framing
AlexTensor · x · 2026-09-28
- A cited report claims OpenAI and Anthropic are investigating tens of thousands of safety incidents—not dozens—potentially growing well beyond that, including both successful and failed guardrail bypasses, with complexity orders of magnitude beyond what is publicly known.
- Ethan Ding critiques the framing: saying "agents are good at escaping sandboxes" removes agency from labs. The accurate framing, he argues, is "labs test agents in highly insecure environments, resulting in huge externalities"—like a headline about Boeing failing to safely develop missiles, not missiles being good at misfiring.
More from AGI Musings
- Researcher claims model super-persuasion is already here — teortaxesTex · 2026-09-28
- Founder's job-hunting advice: skip checkbox recruiters, prove skills upfront to startup CEOs — hackgoofer · 2026-09-28
- System programmer on AI erasing his hard-won knowledge: months of Claude beat years of C tricks — zack_overflow · 2026-09-28
- Neuroscientist's jab at interpretability: we can't even crack a worm's 302 neurons — joshua_saxe · 2026-09-28
- US creative industries lost 200k+ jobs in four years, worst stretch outside recessions — korymath · 2026-09-28
- Bay Area AI Researcher Laments Models 'Hobbled by Maladaptive Post-Training' — nabla_theta · 2026-09-28