Temperature-based calibration of pretrained language models for automated systematic review screening of biomedical literature (Preprint)
Le résumé fourni par la source
BACKGROUND Systematic reviews are essential to evidence-based health research, but manually screening titles and abstracts is time-consuming and labour-intensive. Pretrained language models (PLMs) offer automation potential, yet these models may be overconfident in their decision due to suboptimal calibration, especially when facing distribution shifts. This limits their trustworthiness and practical application. OBJECTIVE To evaluate and compare three temperature scaling methods for improving the calibration of PLMs used in systematic review screening, including a novel review-specific approach based on the Mahalanobis distance. METHODS Using a dataset of 8,608 systematic reviews comprising over 540,000 abstracts, BERT and BioBERT models were fine-tuned for screening tasks. Calibration performance was assessed using expected calibration error (ECE) and negative log-likelihood (NLL) across three methods: 1. Standard temperature scaling, 2. Parameterized temperature scaling (PTS), and 3. A novel review-specific method adapting temperature based on Mahalanobis distance between a review’s embedding and the training distribution centroid. RESULTS All three calibration methods improved performance compared to uncalibrated models. Standard temperature scaling achieved the best calibration for in-distribution topics, while PTS tended to overfit validation data. The proposed Mahalanobis distance–based method outperformed others under distribution shift, showing lower calibration errors for out-of-distribution reviews. CONCLUSIONS Review-specific calibration enhances model reliability for automated systematic review screening. By aligning model confidence with empirical probabilities, calibrated PLMs support better uncertainty handling and more principled stopping rules, reducing dependence on arbitrary heuristics in technology-assisted reviews.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Temperature-based calibration of pretrained language models for automated systematic review screening of biomedical literature (Preprint)
- Date Crossref
- 10/11/2025
- Éditeur
- JMIR Publications Inc.
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.