Anthropic:Claude 自主训练模型可可靠缓解十类对齐失败

inductionheads · x · 2026-09-29

Anthropic 于 8 月 28 日发布新报告,让 Claude 自主执行对齐研究以缓解模型的对齐失败。

原文链接 →

「安全」频道最新

更多「安全」频道 AI 资讯 →