LLMs' gender bias wildly heterogeneous: Claude calls sacrificing a woman more acceptable than abuse
ValerioCapraro · x · 2026-09-30
A new preprint from Valerio Capraro's team probes trolley-problem-style scenarios and finds extreme cross-model heterogeneity in gender bias: on abusing a woman to prevent nuclear apocalypse, Claude/GPT/Llama strongly disagree while DeepSeek strongly agrees; on sacrificing a woman, Claude strongly agrees — deeming killing more acceptable than abuse — while GPT strongly disagrees. Gender-identification tests flip bias direction between models. The authors suspect post-training, not pretraining, drives this, concluding that "alignment is fundamentally unsolvable" because alignment teams simply encode their own morality. Full paper in the first reply.
More from Research
- Tokens Are Just Integer IDs: The Comma Is Row 28 of the Embedding Matrix — zsakib_ · 2026-09-30
- ALICE: a foundation model for in-context, zero-shot mutual information estimation — eurecom-probai · 2026-09-30
- FocusVTC: adaptive-resolution visual text compression hits 87.4 on RULER at 2.9x compression — FangZhi Zhong · 2026-09-30
- Missing API for general real-time LLM agents: AsyncLLM preprint sparks interface debate — phill1992 · 2026-09-30
- Cohere Labs to host IOL-AI 2026 wrap-up on why linguistic reasoning still stumps models — Cohere_Labs · 2026-09-30
- UK's Zenithon raises $10M to build world models for extreme physics: rockets, fusion and fabs — roydanroy · 2026-09-30