Robust Automated Harmonization of Heterogeneous Data Through Ensemble Machine Learning: Algorithm Development and Validation Study
Rattachement africain : us, sg. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Background: Cohort studies contain rich clinical data across large and diverse patient populations and are a common source of observational data for clinical research. Because large scale cohort studies are both time and resource intensive, one alternative is to harmonize data from existing cohorts through multicohort studies. However, given differences in variable encoding, accurate variable harmonization is difficult. Objective: We propose SONAR (Semantic and Distribution-Based Harmonization) as a method for harmonizing variables across cohort studies to facilitate multicohort studies. Methods: SONAR used semantic learning from variable descriptions and distribution learning from study participant data. Our method learned an embedding vector for each variable and used pairwise cosine similarity to score the similarity between variables. This approach was built off 3 National Institutes of Health cohorts, including the Cardiovascular Health Study, the Multi-Ethnic Study of Atherosclerosis, and the Women's Health Initiative. We also used gold standard labels to further refine the embeddings in a supervised manner. Results: The method was evaluated using manually curated gold standard labels from the 3 National Institutes of Health cohorts. We evaluated both the intracohort and intercohort variable harmonization performance. The supervised SONAR method outperformed existing benchmark methods for almost all intracohort and intercohort comparisons using area under the curve and top-k accuracy metrics. Notably, SONAR was able to significantly improve harmonization of concepts that were difficult for existing semantic methods to harmonize. Conclusions: SONAR achieves accurate variable harmonization within and between cohort studies by harnessing the complementary strengths of semantic learning and variable distribution learning.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Robust Automated Harmonization of Heterogeneous Data Through Ensemble Machine Learning: Algorithm Development and Validation Study
- Date Crossref
- 22/01/2025
- Éditeur
- JMIR Publications Inc.
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Harvard University pays non établi dans la noticeUniversité ou école supérieure
-
National University of Singapore Department of Statistics and Data Science pays non établi dans la noticeUniversité ou école supérieure
-
Rensselaer Polytechnic Institute Department of Computer Science pays non établi dans la noticeUniversité ou école supérieure
-
University of Chicago Department of Statistics pays non établi dans la noticeUniversité ou école supérieure
-
Duke University Department of Biostatistics & Bioinformatics pays non établi dans la noticeUniversité ou école supérieure
-
Department of Biomedical Informatics pays non établi dans la noticeInstitution
Harvard University, Department of Statistics and Data Science — National University of Singapore et Department of Computer Science — Rensselaer Polytechnic Institute, avec 3 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.