Code and resources used for assessment, pilot data analysis and data generation in the AIRR-ML-25 (Community Challenge to Benchmark Machine Learning Methods for Adaptive Immune Profiling)
Résumé fourni par la source
Extensive protocol: https://doi.org/10.6084/m9.figshare.31680751Adaptive immune receptor repertoires encode a history of antigen exposure and disease, making them attractive targets for machine-learning (ML)-based diagnostics and therapeutic discovery. However, the absence of standardized benchmarks has hindered objective evaluation of competing methods. We organized the first community challenge to benchmark ML approaches on two central repertoire-analysis tasks: predicting immune state from labelled repertoires (a diagnostic use case) and recovering immune-state-associated receptors (a therapeutic discovery use case), using approximately 75,000 experimentally generated and biologically realistic simulated T-cell receptor repertoires. Across 20 assessed methods, predictive performance remained modest, with the best-performing approaches achieving average ROC AUC values of only 0.70–0.75. Recovery of immune-state-associated receptors proved substantially more challenging, with the top method achieving a Jaccard similarity of approximately 0.15 (range 0-1). By systematically dissecting how performance varied across datasets and study designs, the challenge identifies key methodological limitations, establishes a community benchmark for immune repertoire ML, and provides a roadmap for future advances in ML-based immunodiagnostics and therapeutic discovery.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.