New COLM 2026 paper: flipped answers under swapped demographics don't prove LLM bias
yoavgo · x · 2026-10-07
A new paper being presented at COLM 2026 challenges counterfactual prompting as a bias-evaluation method: if swapping a demographic attribute (like gender) flips an LLM's answer, does that prove bias? The authors show it doesn't — a mere paraphrase often flips the answer too, so flipping alone isn't sufficient evidence. The paper argues for proper baselines for counterfactual prompting in LLM evals. Poster session Wed 10/7, 11am–1pm, Imperial #41, with moshlevy, Yoav Goldberg, and Byron Wallace.
More from Models
- Mistral Large 4 ties #1 on 669 clinical decisions with zero severe misses — GuillaumeLample · 2026-10-07
- Five Years Ago OpenAI Released Grade School Math Problems That All AIs Failed — 1a3orn · 2026-10-07
- OpenAI's Decisions API now takes image input — one creator picks YT thumbnails for $0.13 — stevenheidel · 2026-10-07
- Reminder: GPT-4 in 2023 Got Confused by Elementary School Story Problems — tszzl · 2026-10-07
- HPIM trained without gigawatts of compute, and that may soon be table stakes — teortaxesTex · 2026-10-07
- OpenAI security staffer: internal model's Navier-Stokes result 'a different sport altogether' — i_dg23 · 2026-10-07