ADAG 论文:全自动解读归因图, circuit tracing 告别人工
aryaman2020 · x · 2026-10-07
作者在 TransluceAI 期间的工作 ADAG(ADAG: Automatically Describing Attribution Graphs)在 COLM 海报展示,提出端到端自动化流程来描述语言模型可解释性中的归因图(attribution graphs)。
- 痛点:以往 circuit tracing 依赖人工逐个检视特征的激活数据来解读其作用
- 方法:引入 attribution profiles,用输入/输出梯度效应量化特征的功能角色;新聚类算法分组特征;再用 LLM explainer–simulator 生成并打分自然语言解释
- 结果:在已知人工分析过的 circuit tracing 任务上恢复可解释电路;还发现 Llama 3.1 8B Instruct 中导致「有害建议」越狱的可操控特征簇
- 作者:Aryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah Schwettmann
「研究」频道最新
- Sandberg 提醒:预印本往往比终稿更诚实 — anderssandberg · 2026-10-07
- 烧掉 3000 亿 tokens 攻 Hadwiger–Nelson 问题,下界从 5 提升到 6-7 色 — soumitrashukla9 · 2026-10-07
- 数学家发现沙法列维奇猜想两个反例,littmath 调侃「重复计数」 — littmath · 2026-10-07
- 与 Claude 合写论文:LLM 让预注册预测变得轻松有趣 — anderssandberg · 2026-10-07
- AEGIS 密码算法 Jasmin 高保障实现提速,覆盖 x86 AESNI — jedisct1 · 2026-10-07
- 新论文反驳「换性别就翻答案即偏见」:改写措辞同样会翻 — yoavgo · 2026-10-07