Aller au contenu principal
Accès ouvert déclaré 2026 article

Development and benchmark validation of PubChat for PubMed-grounded multilingual biomedical literature retrieval

0Citations signalées — pas une note de qualité
18Institutions déclarées
5Pays d’affiliation déclarés

Résumé fourni par la source

The rapid growth of biomedical literature has rendered traditional systematic reviews unsustainable. Although large language models (LLMs) offer automation potential, citation fabrication and unreliable evidence discrimination remain critical barriers. Here, using PubChat as a PubMed E-utilities-grounded retrieval framework, we tested whether source-verifiable AI-assisted retrieval could preserve recall, criterion-driven relevance stratification, and multilingual accessibility in systematic-review benchmarking. The system comprises Phase I, hierarchical decomposition of the research question into five relevance levels, and Phase II, multi-round retrieval with embedding-based pre-filtering and three-round LLM verification. Validated against 20 Cochrane systematic reviews (585 ground-truth articles) across eight languages, PubChat was benchmarked against four general LLMs (GPT-5.2-Thinking, Gemini 3.0 Pro, Grok-4.1-Thinking, Qwen3-Max), one search-augmented retrieval tool (Perplexity-Sonar), and three specialized retrieval tools (Elicit, ASTA, and Consensus). PubChat produced no fabricated citations in this benchmark and showed distinct recall–precision profiles across its three operating modes. Among the evaluated configurations, PubChat-Broad achieved the highest observed recall and nDCG, whereas PubChat-Core achieved the highest observed F1- and F2-scores. In a user evaluation of 279 biomedical researchers across 18 countries, PubChat scored above 80/100 for reliability, innovation, efficiency, user experience, and self-reported preference over the evaluated alternatives (78% vs. specialized tools and 72% vs. manual search). An exploratory MIDE case study further illustrates how PubChat-derived corpora can be organized into evidence-traceable research-gap candidates for hypothesis prioritization. PubChat provides a benchmark-validated PubMed-grounded framework for source-faithful biomedical literature retrieval, relevance stratification, and structured evidence organization.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Development and benchmark validation of PubChat for PubMed-grounded multilingual biomedical literature retrieval
Date Crossref
02/09/2026
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Biomedical Text Mining and OntologiesInformation Retrieval and Search BehaviorAcademic Publishing and Open Access

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.