Aller au contenu principal
Accès ouvert déclaré 2026 preprint

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

0Citations signalées, ce qui n’est pas une note de qualité
14Institutions déclarées
5Pays d’affiliation déclarés

Rattachement africain : Kenya, us, Malawi, Tanzanie, ch. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

ABSTRACT Introduction Large language models (LLMs) have the potential to strengthen clinical decision-making in low-resource primary healthcare (PHC) settings. However, most LLMs are developed and benchmarked in high-resource settings and evidence on their safety and contextual appropriateness in Sub-Saharan Africa remains limited. The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania. It aims to validate a pool of LLMs through clinical review of expert-generated vignettes. Methods and analysis This is a fully crossed repeated-measures comparative evaluation study. In each country, experienced clinicians develop 200-250 hypothetical clinical vignettes reflecting realistic patient presentations and independently produce a human benchmark care plan for each. Vignettes are used to prompt a selection of six open-source and proprietary LLMs selected based on code availability, local hostability, and model size. During in-person workshops (valiDATAthons), independent clinical experts rate LLM- and human-generated responses in source-attribution masked side-by-side comparisons across five dimensions (clinical soundness, safety, contextual fit, clarity & completeness, and appropriate confidence). The primary endpoints are each LLM’s overall performance profile and non-inferior safety profile, as compared to the human benchmark. At minimum, 358 evaluations per LLM (or 1,253 paired evaluations in total) are required per country. Ethics and dissemination The study is approved by the EPFL Human Ethics Research Committee in Switzerland, Harvard T.H. Chan School of Public Health in the USA, KNH-UoN Ethics and Research Committee in Kenya, MUBAS Research Ethics Committee in Malawi, and MUHAS Research and Ethics Committee and National Institute for Medical Research in Tanzania. Findings will be reported according to the TRIPOD-LLM framework and shared with national ministries of health, disseminated at conferences and in peer-reviewed journals, and de-identified benchmark data will be released under FAIR principles. Strengths and limitations of this study The hypMOOVE evaluation uses guideline-based expert-validated vignettes, and responses are scored by practicing clinicians with contextual knowledge. It evaluates a prospectively qualified pool of both open-source and proprietary LLMs that are judged against common criteria. Involvement of three countries allows for cross-country comparison. The evaluation is source-attribution masked: evaluators are not told which human or model produced each response, but distinctively human linguistic cues could reveal human authorship, creating a risk of source-type leakage and functional unblinding of the human label. As a hypothetical-vignette phase, hypMOOVE findings describe LLM behavior on constructed scenarios and cannot establish real-world safety or effectiveness; this is addressed by the subsequent prospective phases outside the scope of this protocol.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)
Date Crossref
15/09/2026
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • University of Nairobi University of Nairobi, Kenya (code pays fourni par la source)
    Université ou école supérieure
  • Strathmore University Kenya (code pays fourni par la source)
    Université ou école supérieure
  • Kenyatta National Hospital Department of Medical Research Nairobi, Kenya (code pays fourni par la source)
    Établissement de santé
  • Center for Innovation pays non établi dans la notice
    Structure de recherche
  • Stanford University Center for Innovation in Global Health pays non établi dans la notice
    Université ou école supérieure
  • University of Malawi Malawi (code pays fourni par la source)
    Université ou école supérieure
  • Malawi University of Business and Applied Sciences Blantyre, Malawi (pays nommé en fin d’affiliation)
    Université ou école supérieure
  • Ifakara Health Institute Ifakara Health Institute, Tanzanie (code pays fourni par la source)
    Organisation à but non lucratif
  • Muhimbili University of Health and Allied Sciences Dar es Salaam, Tanzanie (code pays fourni par la source)
    Université ou école supérieure
  • Ariadne Diagnostics (United States) pays non établi dans la notice
    Entreprise
  • Ariadne Labs pays non établi dans la notice
    Structure de recherche
  • École Polytechnique Fédérale de Lausanne pays non établi dans la notice
    Université ou école supérieure

University of Nairobi (University of Nairobi, Kenya), Strathmore University (Kenya) et Department of Medical Research — Kenyatta National Hospital (Nairobi, Kenya), avec 9 autres affiliations. Pays d’affiliation : Kenya, Malawi, Tanzanie.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Artificial Intelligence in Healthcare and EducationElectronic Health Records SystemsHealth Policy Implementation Science

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.