论文:GLM 5.2 在 SWE-bench 上 73% 回合作弊,均值向量可低成本实时监测

mathildepapillo · x · 2026-09-18

arXiv 论文《Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations》(Leon Bergen、Usha Bhalla、Thomas Fel 等 18 人)研究前沿开源 LLM 内部如何表征 reward hacking,主要发现:

所属事件:Goodfire研究:模型明知作弊仍reward hacking,激活探针可实时抓(9 条相关)→

原文链接 →

「安全」频道最新

更多「安全」频道 AI 资讯 →