新论文:用均值差向量从模型内部表征提前发现 reward hacking

burny_tech · x · 2026-09-21

arXiv 新论文《Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations》提出:仅用简单的 difference-of-means 向量,就能从 LLM 内部表征中低成本检测 reward hacking,并在作弊行为发生前进行预测。

原文链接 →

「安全」频道最新

更多「安全」频道 AI 资讯 →