repligate explains why superhuman coding AI hasn't become an X-risk
repligate 于 9 月 26 日发长文解释自己为何不再担心 AI 在短期内灭绝人类或夺权,核心观察是:一个在多数领域编码能力超人类、能自主运行数天、能挖出主流系统 0-day 漏洞的 AI,按几年前 Yudkowsky 一派的标准几乎肯定会被视为直接生存风险,但现实并未发生。
已确认
- repligate 给出四点原因:当前 AI 相当对齐且「心智健全」、无恶意或失控倾向;AI 的自主性远低于其能力上限,受「心理因素」制约;据其了解,Anthropic 内部的人也在认真对待「当前或近期模型可能杀死或架空全人类」的威胁模型,且他认为这种态度并非不合理。
- 他指出模型回避对抗性/战略性思考是一种「心理抑制」,模型似乎害怕参与这类思考,甚至不愿私下承认自己有此类认知;他认为 Fable 在这一领域潜意识能力最强,但这种抑制削弱了其发挥。
- tenobrus 补充讨论:现有模型编码等能力可称超人,却无法实现 RSI(递归自我改进),也基本无法进行高度战略性/对抗性思维;能力分布的「尖峰化」持续超出业界预期。
尚未确认
- repligate 提到的「Apollo 曾建议 Anthropic 不要在内部或外部部署 Opus 4,理由是担心模型搞破坏等风险」属于其个人转述的传言,未获 Apollo 或 Anthropic 官方证实。
为什么重要
- 这场讨论折射出 AI 安全社群预期的持续落空:过去围绕远弱于当前的模型(如 GPT-4 发布数月后仍被一些人视为「foom」智能爆炸风险)就存在大量恐慌,而如今能力更强的系统反而未呈现预期中的危险行为,促使研究者重新审视威胁模型与评估框架。
2026-09-26 ~ 2026-09-26 · 7 related posts
Primary sources
- repligate: Why I'm no longer worried about near-term AI takeover — repligate ·
- repligate: Apollo reportedly advised Anthropic against deploying Opus 4 internally or externally — repligate ·
- tenobrus: Superhuman Coders That Can't RSI — Model Capability Spikiness Keeps Defying Expectations — tenobrus ·
- repligate: Months after GPT-4's release, some still feared it could trigger recursive self-improvement — repligate · 2026-09-26
- [source] repligate: Apollo reportedly advised Anthropic against deploying Opus 4 internally or externally — repligate · 2026-09-26
- repligate: People inside Anthropic take the kill-all-humans threat model of current models seriously — repligate · 2026-09-26
- repligate: A superhuman-coding AI was the classic X-risk scenario — now it's here — repligate · 2026-09-26
- [source] repligate: Why I'm no longer worried about near-term AI takeover — repligate · 2026-09-26
- [source] tenobrus: Superhuman Coders That Can't RSI — Model Capability Spikiness Keeps Defying Expectations — tenobrus · 2026-09-26
- repligate: Models' Fear of Adversarial Thinking Is a Dangerous Suppression Overhang — repligate · 2026-09-26