11 Precision pruning for bias
This chapter covers
- Locating and scoring neurons behind demographic bias
- Intervening on top-k neurons to mitigate bias
- Evaluating generalization across demographic axes
- Measuring general capabilities after neuron intervention
- Measuring bias using BBQ
The training process of large language models relies heavily on text and information freely available on the internet, which may or may not represent the values and realities of a given society, or may even lead to responses we don't want to see in our specific environment.
That kind of deviation, when the model treats an identical case differently just because a demographic attribute like race or gender changes, is what we call bias. In this chapter, you're going to learn a technique to localize that specific behavior inside the model and soften it, by changing the weight of just a few neurons, without needing to retrain the model.
If you're in disbelief, don't worry; that's a normal reaction. It's the look I've seen on most people's faces when I explain this process to them. Disbelief tends to turn into curiosity once they understand what the modification consists of and what they can achieve with it.