Validation of an AI tool for improving MASLD advanced liver fibrosis diagnosis in primary care: a protocol for a feasibility provider-level crossover randomized controlled trial
Résumé fourni par la source
Advanced fibrosis (≥F3) from metabolic dysfunction–associated steatotic liver disease (MASLD) is commonly under-detected in primary care. Given the significant cognitive burden and time constraints already faced by primary care providers, clinical workflow integration is a major hurdle for introducing new diagnostic tools. AI-assisted decision-support tools hold significant promise for overcoming this barrier by seamlessly introducing new diagnostics into the existing practice setting. FibroX, an explainable artificial intelligence (AI) decision-support tool, uses routinely available clinical data to predict the risk of advanced fibrosis. It presents a unique opportunity to close the liver fibrosis ≥F3 under-diagnosis gap in primary care. It accomplishes this by operationalizing dual-threshold triage and providing Shapley additive explanations to support both trust and actionability among clinicians. Real-world application suggests FibroX not only outperforms the standard FIB-4 score but is also prognostic for long-term mortality. However, its feasibility and impact within existing clinical workflows remain untested. As a result, we will investigate these critical parameters in a multi-site, provider-level, randomized crossover simulation pilot. This 12-month pilot will recruit up to 40 primary care clinicians from 4 to 6 clinics. Each clinician will complete two periods after allocation to an AB or a BA sequence where A is FibroX and B is usual care (standard labs ± link to FIB-4). Each provider will diagnose 16 simulated MASLD-risk cases per period separated by a 1-week washout period. Ground truth fibrosis stage for cases will be derived from biopsy or vibration controlled transient elastography-based expert consensus. Primary feasibility will aim for recruitment ≥70%, completion ≥85%, median decision time ≤3.5 min, System Usability Scale ≥70, and AI-Trust neutral-to-positive. Primary effectiveness will be quantified using within-provider diagnostic accuracy for ≥F3. Secondary outcomes will include appropriate referrals, net reclassification improvement, calibration, confidence, cognitive load using NASA Task Load Index, reclassification fairness, and intended downstream testing burden. Implementation outcomes will follow the Reach, Effectiveness, Adoption, Implementation, and Maintenance framework. The McNemar’s test and mixed-effects models will be used for analyses. The authors estimate close to 1150 decisions yielding >80% power to detect 15-point accuracy improvement (i.e., α = 0.05). This pilot will establish feasibility, usability, signal of effectiveness, and implementation readiness for a definitive pragmatic trial. ClinicalTrials.gov. Protocol version: v1.0 (date: October 5, 2025). IRB: Yale University HIC #2000027433. The trial is sponsored by Yale School of Medicine. As the sponsor, Yale School of Medicine holds the ultimate responsibility for the initiation, management, and financing of this clinical investigation, ensuring its conduct adheres to the highest standards of ethics, regulatory compliance, and scientific integrity. The sponsor’s role includes, but is not limited to, regulatory compliance, oversight and quality assurance, safety management, financing, and resource provision.