Accuracy, User Experience, and Implementation Readiness of Consumer Wrist-Worn Wearable Devices Across Six Health Monitoring Domains: Systematic Review and Meta-Analysis
Résumé fourni par la source
Background: Wrist-worn wearable devices have proliferated rapidly, with ownership exceeding 44% in some populations. Contemporary smartwatches now monitor heart rate (HR), atrial fibrillation (AF), blood pressure (BP), oxygen saturation, sleep, and physical activity. Some devices have received regulatory clearance for medical applications, blurring the boundary between consumer wellness tools and medical-grade monitoring equipment. However, evidence on device accuracy and user experience remains fragmented across health domains, populations, and device generations. Objective: To comprehensively evaluate the accuracy, user experience, and implementation considerations of commercially available wrist-worn wearable devices with visual feedback interfaces across six health monitoring domains. Methods: Eight electronic databases were searched from January 2018 to September 2024, supplemented by AI-assisted searching. Independent reviewers screened studies and extracted data. Methodological quality was assessed using Joanna Briggs Institute critical appraisal tools. Quantitative meta-analysis was conducted for HR monitoring (random-effects, mean absolute percentage error, stratified by population and activity context) and AF detection (bivariate random-effects diagnostic test accuracy meta-analysis, stratified by automated versus expert clinician interpretation). The protocol was registered with PROSPERO (CRD42024577994). Results: Fifty studies involving 879,627 participants were included. Apple Watch (33 studies, 66%) and Fitbit (19 studies, 38%) were most frequently evaluated across 45 unique device models. HR monitoring achieved acceptable accuracy (pooled MAPE 4.42%, 95% CI 3.36-5.47%), with a gradient from healthy adults at rest (MAPE 3.82%) to clinical populations during exercise (MAPE 4.93%). Expert clinician interpretation of device electrocardiography (ECG) significantly outperformed automated algorithms for AF detection (sensitivity 0.97 vs 0.87; specificity 0.98 vs 0.96), with automated algorithms missing approximately one in six cases. Inconclusive tracings ranged from 10% to 30% across devices. Energy expenditure (EE) (CV 8.0-34.5%), BP (mean error 6.8-9.4 mmHg), and sleep staging (ICC 0.29-0.64) demonstrated insufficient accuracy for clinical use. Oxygen saturation was acceptable at rest but showed high measurement dropout (30-48%) and near-complete failure during motion. User experience was generally positive under optimal conditions, but device abandonment, technical failures, and measurement dropout limited real-world utility. Extreme data skew was observed, with two studies contributing 99.5% of the total participant pool. Only 8% of studies (4/50) examined device performance by skin tone or race and none were adequately powered to detect differential effects. A further 60% did not report participant race or ethnicity. Only one study explicitly recruited an Indigenous population. Conclusions: Consumer wrist-worn wearables demonstrate sufficient accuracy for wellness tracking across most domains, but performance varies substantially by measurement context, population, and health parameter. HR monitoring and AF detection with mandatory expert ECG review are the most clinically mature applications. EE, BP, and sleep stage classification remain unsuitable for clinical decision-making. Realizing clinical potential requires standardized validation protocols, transparent algorithm reporting, and dedicated research in underrepresented populations and individuals with darker skin tones.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.