Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Pere Martra · hf · 2026-07-31
This work presents Fairness Pruning, a lightweight structural intervention method for managing and mitigating demographic bias in LLMs. Using minimally contrastive prompt pairs and inference-time activation capture, it identifies neurons that react differentially to demographic attributes in GLU architectures. Evaluated on models up to 3B parameters (Llama-3.2 family and Salamandra-2B), results show zeroing these neurons alters responses to demographic variables but causes bidirectional bias destabilization. The intervention is surgical: zeroing at most 40 neurons in Llama-3.2-1B (<0.031% of MLP width) retains 99.49% of reasoning and general knowledge.
More from Research
- Anima Model LoRA Training Struggles with Complex Fine Details — SulpharTriangle · 2026-07-31
- Paper Proposes 'Singleton Attractor' Model to Quantify AGI Capability Thresholds — RecmacfonD · 2026-07-31
- AI-Generated Threat Intelligence Summaries Emerge as Risky New Trend — cyb3rops · 2026-07-31
- New Taxonomy and Observatory for AI 'Scheming' Behaviors Released — S_OhEigeartaigh · 2026-07-31
- DeepSeek V4's fully compressed layers might bottleneck long-context performance — bookwormengr · 2026-07-31
- New Review on Opportunities for Legged Robots by Jonas Frey et al. — ChongZzZhang · 2026-07-31