Bending the Learning Curve for EHR Research via Knowledge-Driven Online Multimodal Automated Phenotyping System
Le résumé fourni par la source
Electronic health records (EHRs) hold great promise for translational research but remain difficult to use at scale because diagnostic codes are noisy, disease-relevant features are hard to identify, high-quality labels are limited and patient-level data sharing is often restricted. We introduce Knowledge-driven Online Multimodal Automated Phenotyping (KOMAP), a unified system for automated, privacy-preserving EHR phenotyping using summary statistics. The system is organized around three original and integrated layers. First, the Multi-source Representation Learning (MultiReL) embedding layer fuses EHR co-occurrence information from multiple health systems with biomedical language-model representations into a shared semantic space. Second, the Online Narrative and Codified feature Engine (ONCE) retrieves and ranks disease-relevant codified and narrative concepts using MultiReL and local EHR support. Third, the KOMAP phenotyping layer trains, validates, and transports disease-specific algorithms from simple summary statistics rather than individual-level patient records. Validation in four healthcare centers demonstrated the ability of KOMAP to support patient subtyping, generate highly accurate phenotyping algorithms, and improve statistical power in downstream phenome-wide association studies (PheWAS). In a PheWAS of a statin-related genetic variant, KOMAP identified significant protective associations for alopecia and seborrheic dermatitis that were not found by standard methods. By combining knowledge-guided feature retrieval with summary-statistics-based phenotyping, the proposed framework reduces privacy, computational, and technical barriers to high-throughput EHR research and supports scalable multi-institutional biomedical discovery.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.