Complexity-driven feature selection for enhancing tuberculosis detection
Résumé fourni par la source
Most existing machine learning approaches for tuberculosis (TB) screening typically utilize large, high-dimensional acoustic feature sets without examining their intrinsic discriminative power. To address this gap, we introduce a complexity-based feature selection approach that evaluates temporal and spectral descriptors using Fisher score (F1), class-distribution overlap (F2), and Shannon entropy (F4). Applied to the CODA-TB dataset (9772 audio recordings from 1105 participants), the proposed method identified 7 highly informative features from the original 26 features, primarily consisting of mel-frequency cepstral coefficients (MFCC) derivatives and spectral-shape measures. The resulting model achieved performance comparable to full-feature baselines while reducing feature dimensionality by 73% and computational cost by up to 14×. Comparative evaluation against four established feature selection techniques, supported by ablation and statistical analyses, confirmed the efficiency and robustness of the complexity-driven strategy, with no statistically significant loss in performance. These findings highlight the potential of lightweight, interpretable, and computationally efficient models for TB cough-based screening in resource-constrained environments.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Complexity-driven feature selection for enhancing tuberculosis detection
- Date Crossref
- 01/12/2026
- Éditeur
- Elsevier BV
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.