New COLM 2026 paper: flipped answers under swapped demographics don't prove LLM bias

yoavgo · x · 2026-10-07

A new paper being presented at COLM 2026 challenges counterfactual prompting as a bias-evaluation method: if swapping a demographic attribute (like gender) flips an LLM's answer, does that prove bias? The authors show it doesn't — a mere paraphrase often flips the answer too, so flipping alone isn't sufficient evidence. The paper argues for proper baselines for counterfactual prompting in LLM evals. Poster session Wed 10/7, 11am–1pm, Imperial #41, with moshlevy, Yoav Goldberg, and Byron Wallace.

Original post →

More from Models

Models channel →