Fine-tuning GPT for gender inclusivity backfires, creating new asymmetric bias, study finds

maier_ak · x · 2026-09-03

A 2024 study by Fulgu and Capraro (University of Milan-Bicocca), covered in Andreas Maier's newsletter, shows that fine-tuning GPT-3.5, GPT-4 and GPT-4o for gender inclusivity creates a new, directionally asymmetric bias.

The key novelty is not that GPT models carry gender bias (well documented), but that this asymmetric bias appears to emerge after inclusivity-oriented fine-tuning.

Original post →

More from Models

Models channel →