Are Your Fairness Metrics Accurate? A Semi-Supervised Approach to Improving Fairness Estimates Under Sample Selection Bias
Rattachement africain : us, es. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
A key challenge impeding the widespread deployment of machine learning is overcoming the impact of statistical biases in the data. Models trained on unrepresentative data can perform worse than anticipated and differentially affect cross-sections of the population. Therefore, evaluating and vetting models based on an appropriate notion of fairness is often indispensable, making accurate estimation of fairness metrics a critical step to safeguard against deployment of unfair algorithms. It is often assumed that a fairness metric computed from the observed data is accurate. However, in presence of selection bias, also referred to as distributional shifts, fairness metric estimates too can have systematic application-specific errors. In this work we demonstrate this phenomenon and, relying on access to an unbiased unlabeled data, derive a semi-supervised approach to mitigate estimation errors emerging from the biased labeled data. Specifically, we introduce a novel selection bias model called ''sub-class-conditional invariance'' (SCC-invariance), that offers a flexible framework to effectively capture distributional shifts in the real-world data, particularly compared to traditional models such as label shift and covariate shift. Assuming a finite Gaussian mixture form for each class-conditional distribution, we then derive an Expectation-Maximization algorithm to estimate model parameters and correction weights necessary for computing unbiased estimates. We focus on three widely used fairness metrics--equal opportunity, predictive equality, and predictive parity--and demonstrate the effectiveness of our approach in improving their estimates on synthetic data. Finally, we apply our bias mitigation approach to clinical genetics and study the fairness of pathogenicity predictors across ancestral groups.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Are Your Fairness Metrics Accurate? A Semi-Supervised Approach to Improving Fairness Estimates Under Sample Selection Bias
- Date Crossref
- 03/08/2025
- Éditeur
- ACM
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.