Study finds all 13 major LLMs flip truth judgments on speaker gender, up to 23.6% of statements
anthara_ai · x · 2026-09-03
An experiment across 13 popular LLMs found every model shows statistically significant negative sentiment toward men.
Method: researchers altered only the speaker's gender presentation (neutral/male/female) and measured whether truth judgments changed.
Key findings:
- 10%–35% of statements received completely inconsistent truth labels solely due to gender presentation
- Flip rates hit up to 23.6% when comparing male vs. female variants
- Every tested model exhibited gender sensitivity
The results highlight hidden consistency risks for automated systems that rely on LLM judgments.
More from Models
- Claude is down for many users amid widespread outage — Polymarket · 2026-09-03
- Claude Mythos 5.1, Fable 5.1 and Opus 5 hit elevated errors, Anthropic investigating — ClaudeAI-mod-bot · 2026-09-03
- Two-Author Model Tech Report Praised as Dense: Pretraining to Downstream — antoine_chaffin · 2026-09-03
- IFM open-sources K2 Horizon models from 0.9B to 375B with training code and data recipes — testingcatalog · 2026-09-03
- Leaked hints suggest the next release will be v2.5 — koltregaskes · 2026-09-03
- MBZUAI releases K2 Horizon: six fully open models from 0.9B to 375B with training code and data — kimmonismus · 2026-09-03