BiasGym, an injection-based LLM bias analysis and removal framework, accepted at EMNLP 2026

IAugenstein · x · 2026-09-02

"BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection" from Isabelle Augenstein's group was accepted to EMNLP 2026 Findings. The framework injects fictional tokens to surface biases in LLMs, enabling systematic analysis and removal. Paper and code are public.

Original post →

More from Safety

Safety channel →