Innovations In Machine Assessment Of Replicability
Le résumé fourni par la source
Automated methods for the assessment of replicability of scientific claims offer a scalable complement to replication studies and traditional peer review. Drawing on a large dataset of claims, human judgments, and a limited set of replication outcomes, we developed and evaluated three distinct artificial intelligence systems designed to predict human expert assessments of replicability using diverse methodologies—including synthetic prediction markets, interpretable feature-based modeling, knowledge graph reasoning, and semantic parsing with argument structures. While these systems achieved modest calibration to human judgment distributions, they failed to discriminate between replicable and non-replicable claims. Our findings suggest that while machine assessments of research replicability may complement human reasoning, their current performance limitations and opportunities for bias demand careful evaluation before real-world application.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.