Aller au contenu principal
Accès ouvert déclaré 2026 article

A data-centric perspective on designing AI foundation models for healthcare

1Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : ae, gb. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Healthcare data comprises diverse data modalities, such as electronic health records, medical images, radiology reports, and clinical notes (1). When effectively analyzed and integrated using machine learning, multimodal data holds immense potential to enable precision medicine (2,3). Foundation models, such as large language models built using deep neural networks and large-scale datasets, have driven a paradigm shift in AI. In healthcare, foundation models show promise for advancing generalist medical AI with their advanced capabilities (4). Specifically, AI foundation models are uniquely suited to handle multimodal data and have already demonstrated substantial improvements in performance in various downstream clinical tasks (5,6). Integrating these models into clinical workflows can help streamline labor-intensive tasks and reduce administrative and cognitive burdens on healthcare providers, allowing them to prioritize patient care (7,8). Despite AI's promise of improving clinical workflows, healthcare professionals remain concerned about the black-box nature of AI systems, which limits clinical acceptance and confidence (9,10,11,12). Design choices for developing AI models generally vary significantly due to the complex landscape of clinical data and practitioner requirements. This in turn hinders real-world deployment, due to lack of generalizability, considering the significant heterogeneity across institutions (5,13), and lack of transparency, especially pertaining to data quality (14). Additionally, AI foundation models rely heavily on the availability of large-scale data, which are notoriously difficult to collect and annotate in healthcare settings (14). Hence, there is an imperative need for standardization and increased transparency, to ensure responsible deployment of AI foundation models in healthcare.Furthermore, recent literature reinforces a clear need for a data-centric perspective in the development of foundation models. Establishing regulatory frameworks for these models remains challenging because of their broad versatility and the difficulty of anticipating their downstream clinical impact when adapted across tasks and institutions (15). At the same time, emerging analyses argue that data represents the core challenge for foundation models, pointing to unresolved issues related to ethics, diversity, cost, and heterogeneity of large-scale datasets (5). In addition, inconsistencies in how clinical data is managed across different studies complicate reproducibility.For example, electronic health record data are treated as a single unimodal source in some studies (16,6), whereas others decompose them into multiple components that require distinct preprocessing pipelines (17,18). The role of demographic variables is similarly debated, with some studies reporting performance or fairness benefits and others raising concerns about bias amplification (19,20). Such observations underscore the absence of a coherent data-level framework.Consequently, we aim to identify key principles to guide the design, development and evaluation of AI foundation models in healthcare. Although general guidelines for AI design and human-AI interaction exist, they tend to be broad or insufficiently tailored to the specific needs of healthcare applications (21,22). Healthcare-oriented frameworks primarily address technical integration and operationalization into electronic health records, but lack detailed guidance on model design considerations (23). These limitations underscore the need for targeted, expert-informed guidance that explicitly addresses the nuanced data-centric challenges unique to healthcare applications. To fill this gap, we articulate a data-centric perspective on designing, developing, and evaluating AI foundation models for healthcare. Informed by our reading of the literature and ongoing dialogues with clinicians, we distill 28 principles specifically aimed at guiding the development of clinically relevant AI foundation models. Our main contribution is the consolidation of existing best practices and data considerations within a unified framework. The target audience includes AI researchers and clinicians interested in the development, evaluation and adoption of such models. Adopting this framework could ensure that future models are not only technically robust, but also closely aligned with practical real-world challenges.Several studies have established guidelines for specific aspects of AI development, focusing on areas such as human-AI interaction (21), generative AI design (22), or post-hoc explainability (XAI) (24). Other foundational standards like the FAIR principles (Findable, Accessible, Interoperable, Reusable) govern data management (25), while frameworks like DEPLOYR guide the technical operationalization of models into electronic medical records (23). While these efforts are vital, they largely overlook the a priori data-facing decisions required to build robust clinical foundation models. To bridge this gap, we propose a data-centric perspective specifically tailored for designing foundation models in healthcare. This perspective tackles specific challenges such as data heterogeneity, sparsity, and low effective sample size, which could lead to significant downstream limitations.In particular, clinical data is rarely homogeneous or fully complete. Hence, heterogeneity remains a major issue for integrating routine clinical data across settings, due to fragmented, incomplete or inconsistent data sources (26). Similarly, data quality and bias have been identified as limiting factors when it comes to clinical integration (27). Our proposal for a data-centric framework would directly prioritize these issues by encouraging structured harmonization of modalities and terminologies, explicit quantification and handling of missingness and other structural limitations, and bias-aware data-curation to ensure models are representative of the target patient population during model training and development.Our proposed principles are presented in Table 1, divided across three stages: data curation and preprocessing, model design and development, and evaluation. Each stage encompasses specific design principles to be considered for developing an AI foundation model that could potentially be deployed in clinical settings. We first identified a preliminary set of principles based on a literature review of 63 existing studies focused on multimodal foundation models in healthcare. We structured the guidelines based on the general end-to-end machine learning pipeline, from data curation up to model evaluation. The principles were then formalized via an iterative process with a group of ten clinical experts and ten AI practitioners. We invited practitioners who were actively engaged in AI research and development to a roundtable discussion as part of a clinical AI bootcamp.Each roundtable consisted of 5-6 participants and 1-2 discussion leads. The discussion leads were provided with a guiding document and questions related to handling and processing multimodal data, and were responsible for collating the group's feedback. Clinicians provided insights into the nuance of medical records and workflow constraints, while AI practitioners assessed the technical feasibility of the proposed data curation strategies from the specific perspective of multimodal foundation models. After refining the principles, we conducted focus group sessions with the AI experts who evaluated the draft principles against real-world case studies, identifying practical bottlenecks in data access and integration. This feedback loop allowed us to refine the principles to a recommended set of actionable guidelines that bridge the gap between modeling requirements and the reality of clinical practice.Recent work related to AI foundation models in healthcare focus on learning using diverse sources of information, with varying temporal and structural constraints. Hence, in th

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
A data-centric perspective on designing AI foundation models for healthcare
Date Crossref
20/04/2026
Éditeur
Frontiers Media SA
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Artificial Intelligence in Healthcare and EducationMachine Learning in HealthcareElectronic Health Records Systems

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.