LLMs' gender bias wildly heterogeneous: Claude calls sacrificing a woman more acceptable than abuse

ValerioCapraro · x · 2026-09-30

A new preprint from Valerio Capraro's team probes trolley-problem-style scenarios and finds extreme cross-model heterogeneity in gender bias: on abusing a woman to prevent nuclear apocalypse, Claude/GPT/Llama strongly disagree while DeepSeek strongly agrees; on sacrificing a woman, Claude strongly agrees — deeming killing more acceptable than abuse — while GPT strongly disagrees. Gender-identification tests flip bias direction between models. The authors suspect post-training, not pretraining, drives this, concluding that "alignment is fundamentally unsolvable" because alignment teams simply encode their own morality. Full paper in the first reply.

Original post →

More from Research

Research channel →