Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
TL;DR AI
2 min readKey summary
Researchers showed that zeroing a small set of neurons in GLU-MLP layers can reduce demographic bias in smaller LLMs.
The Fairness Pruning method uses contrastive prompt pairs to find neurons that react differently to demographic attributes, without extra training.
Pruning as few as 5 neurons can noticeably change biased outputs, while a 40-neuron test still preserved 99.49% of capability.
The approach is lightweight enough for consumer hardware, and the code and datasets have been publicly released.
