BiasGym 框架被 EMNLP 2026 接收,注入法定位并消除 LLM 偏见

IAugenstein · x · 2026-09-02

Isabelle Augenstein 组的论文《BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection》被 EMNLP 2026 Findings 接收。该框架提出通过向模型注入虚构 token 来「诱导」偏见显形,从而系统地分析和移除大模型中的偏见,论文与代码均已公开。

原文链接 →

「安全」频道最新

更多「安全」频道 AI 资讯 →