C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
Rattachement africain : sg, cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Bias in Large Language Models (LLMs) poses significant risks to trustworthiness, manifesting primarily as stereotypical biases (e.g., gender or racial stereotypes) and structural biases (e.g., lexical overlap or position preferences).However, prior paradigms typically address these in isolation, often mitigating one at the expense of exacerbating the other.To address this, we conduct a systematic exploration of these reasoning failures and identify a primary inducement: the latent spurious feature correlations within the input that drive these erroneous reasoning shortcuts.Driven by these findings, we introduce Causal-Contrastive Preference Optimization (C2PO), a unified alignment framework designed to tackle these specific failures by simultaneously discovering and suppressing these correlations directly within the optimization process.Specifically, C2PO leverages causal counterfactual signals to isolate biasinducing features from valid reasoning paths, and employs a fairness-sensitive preference update mechanism to dynamically evaluate logitlevel contributions and suppress shortcut features.Extensive experiments across multiple benchmarks covering stereotypical bias (BBQ, Unqover), structural bias (MNLI, HANS, Chatbot, MT-Bench), out-of-domain fairness (Stere-oSet, WinoBias), and general utility (MMLU, GSM8K) demonstrate that C2PO effectively mitigates stereotypical and structural biases while preserving robust general reasoning capabilities.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
- Date Crossref
- 01/01/2026
- Éditeur
- Association for Computational Linguistics
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Nanyang Technological University pays non établi dans la noticeUniversité ou école supérieure
-
University of Jinan pays non établi dans la noticeUniversité ou école supérieure
-
Shanghai Key Laboratory of Trustworthy Computing pays non établi dans la noticeStructure de recherche
-
Engineering Research Center of Trustworthy AI (Ministry of Education) pays non établi dans la noticeStructure de recherche
-
Jinan University pays non établi dans la noticeUniversité ou école supérieure
-
Guangxi Key Laboratory of Trusted Software pays non établi dans la noticeStructure de recherche
Nanyang Technological University, University of Jinan et Shanghai Key Laboratory of Trustworthy Computing, avec 3 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.