GPT-OSS Openness vs Safety: Aidan Clark Responds to Critics

A multi-faceted debate has erupted among researchers like Aidan Clark regarding the open-source and safety strategies of GPT-OSS. The core disagreement lies in the actual risks of open models and whether the safety team's restrictions are justified. This discussion touches upon technical risk assessment and reflects the underlying tensions among the open-source community, AI companies, and safety researchers.

Clash of Viewpoints

Aidan Clark expressed disappointment that the focus of the open-source discussion has shifted from "safety" to "sovereignty." He noted that GPT-OSS was intended to benefit everyone, but the team delayed its launch specifically to ensure the built-in safety guardrails were robust enough. He emphasized that the risk threshold for frontier models is difficult to judge in advance, making iterative deployment strategies essential; one cannot simply claim that early safety concerns were unwarranted just because things look safe in hindsight. Using GPT-2 as an example, he argued that while delaying the release might seem unwise in retrospect, those safety concerns were entirely reasonable given the incomplete information at the time.

Controversy and Skepticism

Critics have pushed back. Some questioned why GPT-OSS couldn't simply be released as an uncensored version if it posed no safety risks, bluntly stating that "AI safety" is often just a cover for corporations avoiding reputational and legal risks. However, researchers like David Manheim argued that open models do present legitimate reasons for caution, such as targeted abuse. Clark responded by noting that both the "corporate shills" and the "open-source die-hards" carry their own biases, as this is not an ordinary software debate. He also suggested that there is actually less to worry about with the OSS release now, implying that the team's real reason for holding back might simply be other pending priorities.

2026-07-21 ~ 2026-07-21 · 9 related posts