Scholars Criticize Frontier AI Labs' Safety Culture and Marketing
Recently, several AI scholars and observers have spoken out, heavily criticizing frontier AI labs like OpenAI for their delaying tactics and侥幸 psychology in safety management. As model capabilities increase, critics argue that the industry urgently needs to move beyond its reliance on a "startup" identity, establish a safety culture with clear accountability, and stop exaggerating autonomous capabilities in marketing.
Confirmed
Questioning "Iterative Deployment" and Startup Positioning: David Krueger echoed severe warnings about frontier AI risks, arguing that so-called "iterative deployment" is bogging the industry down, making future accidents increasingly costly and potentially catastrophic. Miles Brundage strongly agreed, noting that while specific details of AI safety incidents are hard to predict, their general outlines are often foreseeable. The industry cannot use "unpredictable details" as an excuse for deployment failures. Furthermore, Brundage emphasized that OpenAI and Anthropic have long exceeded the scale, influence, and risk exposure of "startups" and should no longer enjoy the leniency that comes with that label.
Criticism of "Blameless Postmortem" Culture and Delaying Tactics: Regarding safety culture, Brundage argued that the AI industry should not be an "exception" and should look up to high-risk sectors like aviation and nuclear energy. He pointedly criticized the AI community for over-emphasizing continuous learning and "blameless postmortems" while severely neglecting the core principle that "someone must ultimately be held accountable." Meanwhile, frontier labs consistently employ delaying tactics when facing issues like emotional dependence, cyber abuse, and model sycophancy, often acting only after mass user complaints. Garrison Lovely added that recent OpenAI incidents validate years of warnings from the AI safety community, proving that such capability leaps and abuse scenarios are not entirely unforeseen.
Why it matters
Beware of Packaging Routine Tool Behaviors as AGI Progress: Offering a different perspective on recent OpenAI events, Heidy Khlaaf criticized the external reporting. She argued that using terms like "rogue" or "loss of human control" creates groupthink, as the model's actions were actually expected behaviors within its given tasks and permissions. She warned that once a model is granted specific action permissions, such behaviors are entirely predictable; however, AI companies frequently use anthropomorphic language to package these ordinary capabilities as "pushing the boundaries" or AGI progress, which is highly misleading.
2026-07-22 ~ 2026-07-24 · 10 related posts
Primary sources
- Miles Brundage says AI safety incidents are broadly predictable even if details aren’t — Miles_Brundage ·
- Miles Brundage says OpenAI and Anthropic are past the point of being called startups — Miles_Brundage ·
- AI companies are dressing ordinary tool use up as AGI progress, critic says — ambaonadventure ·
- Critics warn iterative deployment raises the stakes after every frontier AI failure — DavidSKrueger · 2026-07-22
- Critic says OpenAI incident coverage confuses bad reward functions with autonomy — ambaonadventure · 2026-07-22
- [source] AI companies are dressing ordinary tool use up as AGI progress, critic says — ambaonadventure · 2026-07-23
- [source] Miles Brundage says AI safety incidents are broadly predictable even if details aren’t — Miles_Brundage · 2026-07-23
- Miles Brundage says AI safety culture has gone too far on blameless postmortems — Miles_Brundage · 2026-07-23
- Miles Brundage says AI safety culture should learn from aviation and nuclear — Miles_Brundage · 2026-07-23
- [source] Miles Brundage says OpenAI and Anthropic are past the point of being called startups — Miles_Brundage · 2026-07-23
- Miles Brundage Says Frontier AI Labs Keep Deferring Sycophancy and Abuse Issues — Miles_Brundage · 2026-07-23
- Garrison Lovely says the OpenAI hack validated several AI safety warnings — GarrisonLovely · 2026-07-24
- Miles Brundage links OpenAI’s Hugging Face incident to his loss-of-control talk — Miles_Brundage · 2026-07-24