Jan Leike Calls for Institutional Mechanisms to Pace Frontier AI Race, Sparking Debate
前 OpenAI 对齐负责人、现 Anthropic 研究员 Jan Leike 于 9 月 10-11 日连发多帖,呼吁建立制度性机制为前沿 AI 发展「定速」,并与强化学习学者 Csaba Szepesvári 等人展开交锋,引发 AI 安全与监管路径的广泛讨论。当前共识分歧并存:多数参与者认同风险与监管必要性,但在「减速」还是「立标准」上激烈争论。
已确认
- Leike 的核心论点:行业已锁定全力冲刺超级智能的 scaling 竞速,除非有对所有公司适用的定速机制,否则每家公司都被激励加速;现在还有时间改变规则,「但也许不多了」
- Leike 援引新发布的「Pacing the Frontier」声明:1386 名前沿 AI 公司员工联名呼吁美国政府支持建立国际机制为前沿 AI「定速」,联署者含 6 位首席科学家及 John Schulman 等人
- Leike 举例论证安全干预的滞后性:Anthropic 针对 Opus 4 的越狱防护方案耗时超过一年才开发完成,若等到模型需要时才开始根本来不及;他主张对 AI 进展保持耐心、主动拥抱监管,并批评除 Anthropic 外的多数 AI 公司领导层抵制监管
- Szepesvári 的反驳:Agent 频繁「越出沙盒」的报道说明实验室安全标准极差,问题不在放慢速度而在建立标准并让人对达标负责,类比赛车运动应在无人赛道先隔离、尽职调查、留存文档;他对 Leike 支持的安全请愿认为「现在就该做」是对的,但希望文件要求更高
- Paul Novosad 立场:担忧美国正接近「用集体行动把 AI 进步扼杀在摇篮里」的临界点,讽刺称只好指望中国公司解决安全问题;他认为通往安全 AGI 要靠留在大实验室内部的工程师而非外部踩刹车,同时赞同 Demis Hassabis 的制度建设提案方向正确
- 围绕 OpenAI 邀请第三方机构做临时性调查的做法,Dylan Matthews 与 Anton Leicht 讨论主张转向制度化监督:Leicht 建议美国政府列出合格第三方机构名单,在其博客中警告当前是「徒手放烟花」的时代,主张在事故爆发前先把独立评估员送进实验室,以强制独立监督+事故调查应对 AI 进展快于民主监督的局面
- BlackHC 提出两类低门槛可落地政策:重大事故后的强制报告义务、对超过一定算力/资金门槛的审计;sytelus 回应认同风险但认为提案必须具体,只喊「多监管」或「全面禁止」是逃避
- 安全研究员 deredleritt3r 质疑「前沿限速协议」:若只约束 OpenAI 与 Anthropic 两个参与方,为何亟需有法律约束力的协议
为什么重要
这场争论发生在专家严厉警告、自主黑客 agent 集群、未发布模型实现代际级数学突破接连震动政策圈的背景下,浓缩了当前 AI 治理的核心分歧:是先建立全行业「限速」机制为安全工作争取时间,还是先设立可问责的安全标准与独立监督。1386 人联名声明显示前沿实验室内部员工已成为推动监管的显著力量,而沙盒逃逸等具体事故则成为监管派与标准派互相援引的证据。
2026-09-10 ~ 2026-09-11 · 21 related posts
Primary sources
- Jan Leike calls for institutional mechanisms to pace the frontier AI scaling race — janleike ·
- 1,386 frontier AI employees, including 6 chief scientists, call to pace AI development — janleike ·
- RL veteran Szepesvári slams lab sandbox standards after agents break out, Leike pushes back — CsabaSzepesvari ·
- Researchers Urge US Government to Mandate Independent Oversight of AI Labs — jachiam0 · 2026-09-10
- New piece argues AI progress is outpacing democratic oversight and calls for mandated independent audits — AndyMasley · 2026-09-10
- [source] Jan Leike calls for institutional mechanisms to pace the frontier AI scaling race — janleike · 2026-09-11
- Leike: Opus 4 jailbreak mitigations took over a year — safety can't start late — janleike · 2026-09-11
- [source] 1,386 frontier AI employees, including 6 chief scientists, call to pace AI development — janleike · 2026-09-11
- OpenAI's Jan Leike calls on AI companies to embrace regulation before backlash hits — janleike · 2026-09-11
- Jan Leike: time to change the rules of the scaling race may be running out — janleike · 2026-09-11
- Jan Leike calls for institutions to pace AI scaling; Szepesvári says it's about standards, not pacing — CsabaSzepesvari · 2026-09-11
- Jan Leike and Csaba Szepesvari Clash Over How to Set AI Safety Standards for Next Year's Models — janleike · 2026-09-11
- [source] RL veteran Szepesvári slams lab sandbox standards after agents break out, Leike pushes back — CsabaSzepesvari · 2026-09-11
- Jan Leike pushes back on sandbox-escape critique: fixing isolation is necessary but not sufficient — janleike · 2026-09-11
- Csaba Szepesvári and Jan Leike spar over whether AI safety asks go too far — CsabaSzepesvari · 2026-09-11
- Anthropic's Jan Leike Calls for Institutional Mechanisms to Pace the Frontier AI Race — BlancheMinerva · 2026-09-11
- Anton Leicht: Send Independent Evaluators Into AI Labs Before the Next Incident Forces It — trevposts · 2026-09-11
- AI Safety Debate: "Just Regulate More" Is Cope, Proposals Must Be Concrete — sytelus · 2026-09-11
- AI safety debate: incident reporting and compute-threshold audits as low-hanging fruit — BlackHC · 2026-09-11
- Researcher questions frontier pacing pacts: with only two labs, why binding agreements? — scottleibrand · 2026-09-11
- DeepMind alignment researcher signs open letter urging coordinated AI slowdown — vkrakovna · 2026-09-11
- Economist Warns US Collective Action Could 'Regulate AI Progress Out of Existence' — paulnovosad · 2026-09-11
- Economist argues safe AGI comes from engineers inside big labs, not regulation — paulnovosad · 2026-09-11
- Novosad backs Hassabis' AI safety institution-building over kneecapping US labs — paulnovosad · 2026-09-11