MedMI-Bench: Pre-Registered Analysis Plan for Supplementary Experiments
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Pre-registered analysis plan for three supplementary experiments extending the MedMI-Bench physician audit of LLM-generated medication decision benchmarks derived from MIMIC-IV. This document commits to hypotheses, sample sizes, statistical tests, and decision rules before data collection to protect the analysis from post-hoc selection effects. The three experiments are: 1. A position-shuffling causal test on 2,000 consensus-accepted probes. The original 34,201-probe pool analysis showed an observational position bias (13.8% consensus acceptance for position A vs 16.0% for position D, P=.002). This experiment shuffles the correct-answer position while holding clinical content constant and re-runs all three validators (GPT-4o, Claude Sonnet 4.6, Gemini 2.5 Flash) to determine whether the effect is causally driven by position or confounded with content. 2. A 100-probe physician audit of the consensus-rejected pool (n=16,386), matching the 100-probe accepted-pool audit from the main paper, stratified by task type and department. This enables bidirectional characterization of validator error (false-acceptance rate vs false-rejection rate) using matched sample sizes and equivalent confidence interval widths. 3. A cross-validator disagreement analysis on the full 34,201-probe pool, computing pairwise Cohen's kappa between validator pairs, sole-dissenter identity distributions, and a correlated-error test comparing observed joint-error counts against the product of individual error rates under the statistical independence null. Decision rules for paper integration, deviation policies, and null-result reporting commitments are specified in the attached document. This pre-registration accompanies the main MedMI-Bench paper targeting JMIR AI and will be cited in the paper's Methods section with its Zenodo DOI. Author: Aman Sharma (ORCID: 0009-0005-5107-4485)Contact: Aman_sharma007@yahoo.comMain paper target venue: JMIR AIPre-registration timestamp: April 2026 (before experiment execution)
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.