Fine-tuning GPT for gender inclusivity backfires, creating new asymmetric bias, study finds
maier_ak · x · 2026-09-03
A 2024 study by Fulgu and Capraro (University of Milan-Bicocca), covered in Andreas Maier's newsletter, shows that fine-tuning GPT-3.5, GPT-4 and GPT-4o for gender inclusivity creates a new, directionally asymmetric bias.
- Phrase-attribution tests: models systematically label masculine-coded sentences as written by women. Inclusivity index: GPT-3.5 scored 0.535 for masculine-coded vs 0.050 for feminine-coded; GPT-4 0.312 vs 0.043; GPT-4o 0.250 vs 0.027 — all highly significant (p < .001).
- Moral-dilemma tests: GPT-4 unanimously rated harassing a woman to avert disaster as 1, but harassing a man averaged 3.3; removing gender cues closed the gap, indicating implicit bias.
The key novelty is not that GPT models carry gender bias (well documented), but that this asymmetric bias appears to emerge after inclusivity-oriented fine-tuning.
More from Models
- Ex-Cursor RL Lead Joins Meta's Reasoning Team as Muse Spark 1.3 Ships — ananyaku · 2026-09-03
- Devs warn using Gemini outside official surfaces can get your entire Google account banned — GlenBradley · 2026-09-03
- ChatGPT nails a backgammon dice probability problem at 47% odds — ivan_bezdomny · 2026-09-03
- H3 Acceleration Arena needs 1,700 more votes to crown the best turbo LoRA — Obvious_Set5239 · 2026-09-03
- muse in contributor mode is the cheapest high-end model, open or closed — philfung · 2026-09-03
- Latent Space weekly: Meta's Muse Spark 1.3 hits #3 globally, Gemini 3.8 Flash Cyber launches — Latent Space · 2026-09-03