Aller au contenu principal
Accès ouvert déclaré 2026 article

Benchmarking generative AI tools for interpretation of the WHO TB mutation catalogue

2Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : ch, us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

is a crucial knowledgebase and tool for clinical interpretation of mutations associated with drug-resistant TB. However, the document's complexity and size pose challenges for many users. This study evaluated the potential of generative artificial intelligence (AI) models to facilitate natural language user interaction with the catalogue. This was a benchmarking study, not a clinical usability trial. Four prominent AI models-Google Gemini 2.5 Pro, OpenAI ChatGPT 4.1, Perplexity AI, and DeepSeek R1-were assessed through general test questions, mutation search and retrieval tasks using both full catalogue queries and antibiotic-specific tables, and the application of additional grading rules to score novel mutations. Performance was measured based on accuracy, completeness, clarity, source citation, and the presence of hallucinations. Google Gemini 2.5 Pro consistently demonstrated superior performance in accuracy, completeness, and avoidance of hallucinations across most evaluations, especially in general queries and large dataset searches. DeepSeek R1 excelled in applying grading rules to novel mutations and showed high accuracy in focused datasets, but exhibited some hallucinations. ChatGPT 4.1 was strong in clarity but lacked proper citations, and Perplexity AI showed variable performance with a higher frequency of hallucinations. The findings highlight the potential of AI tools to enhance the accessibility of complex knowledgebases like the WHO Mutation Catalogue, while emphasizing the need for rigorous benchmarking. While no model is yet suitable for direct clinical use, the results suggest that with further development, models like Google Gemini 2.5 Pro could form the basis of a custom AI agent to assist users in navigating this critical resource, ultimately contributing to improved TB control efforts.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Benchmarking generative AI tools for interpretation of the WHO TB mutation catalogue
Date Crossref
09/02/2026
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Artificial Intelligence in Healthcare and EducationGenomics and Rare DiseasesMachine Learning in Healthcare

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.