Aller au contenu principal
Accès ouvert déclaré 2025 conference-paper

ASTRID - An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems

1Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : gr. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Large Language Models (LLMs) have shown impressive potential in clinical question answering (QA), with Retrieval Augmented Generation (RAG) emerging as a leading approach for ensuring the factual accuracy of model responses.However, current automated RAG metrics perform poorly in clinical and conversational use cases.Using clinical human evaluations of responses is expensive, unscalable, and not conducive to the continuous iterative development of RAG systems.To address these challenges, we introduce ASTRID -an Automated and Scalable TRIaD for evaluating clinical QA systems leveraging RAG -consisting of three metrics: Context Relevance (CR), Refusal Accuracy (RA), and Conversational Faithfulness (CF).Our novel evaluation metric, CF, is designed to better capture the faithfulness of a model's response to the knowledge base without penalizing conversational elements.Additionally, our metric RA captures the refusal to address questions outside of the system's scope of practice.To validate our triad, we curate a dataset of over 200 real-world patient questions posed to an LLM-based QA agent during surgical follow-up for cataract surgery -the highest volume operation in the worldaugmented with clinician-selected questions for emergency, and clinical and non-clinical out-of-domain scenarios.We demonstrate that CF predicts human ratings of faithfulness more accurately than existing definitions in conversational settings.Furthermore, using eight different LLMs, we demonstrate that the three metrics can closely agree with human evaluations, highlighting the potential of these metrics for use in LLM-driven automated evaluation pipelines.Finally, we show that evaluation using our triad of CF, RA, and CR exhibits alignment with clinician assessment for inappropriate, harmful, or unhelpful responses.We also publish the prompts and datasets for these experiments, providing valuable resources for further research and development.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
ASTRID - An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems
Date Crossref
01/01/2025
Éditeur
Association for Computational Linguistics
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Topic Modeling

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.