A new preprint on arXiv questions a common assumption in algorithmic fairness: that stripping demographic information from a model's internal representations is the best way to prevent bias. The authors note that encoded demographic information is widely treated as a risk factor, with invariance to demographics often touted as the ideal condition.

However, the paper's central claim, reflected in its title, is that this approach is not sufficient — and can even backfire. Removing demographic information may create new forms of bias, the authors argue, suggesting that fairness requires more than simply making representations invariant.

The abstract is truncated in the source, so the specific mechanisms or experiments behind this claim are not detailed here. As a preprint, the work has not yet undergone peer review.