S1034 Automated Interpretation of EndoFLIP Metrics Per Dallas Consensus Using Text- and Image-Based LLM Prompting
Le résumé fourni par la source
Introduction: The functional lumen imaging probe (EndoFLIP) is a novel tool for evaluating esophageal diseases. Recent guidelines (Dallas Consensus) provide structured algorithms for interpreting EndoFLIP data, but applying them can be challenging, especially in centers without dedicated esophageal teams. Large language models (LLMs), capable of understanding unstructured natural language, have shown utility in similar clinical reasoning tasks. More recently, vision-language models have emerged as a more convenient option, able to interpret image-based prompts with minimal or no textual explanation—potentially streamlining algorithm-based diagnostics. This study aims to evaluate the performance of text-based versus image-based prompting for interpreting EndoFLIP data using consensus algorithms. Methods: We created 21 simulated cases using EndoFLIP metrics across 7 diagnostic categories, modeled after real patients from our esophageal disease center, which were evaluated by 5 LLMs of varying sizes. Image-based prompts used visual representations of the consensus algorithms and tables, while text-based prompts included detailed manual descriptions of diagnostic criteria. Prompts were iteratively refined to optimize diagnostic accuracy and ensure model comprehension. Model outputs were compared to expert diagnosis, with percent agreement and diagnostic accuracy calculated. Results: LLaMA-4-Maverick achieved the highest agreement with 81% accuracy for both text- and image-based prompts. In contrast, smaller models like gemma-4b showed the lowest image-based agreement (28.6%) compared to 57.1% with text input. Although differences between prompt modalities were not statistically significant due to smaller sample size, trends favored text input. Conclusion: EndoFLIP interpretation using LLMs is feasible. While vision-based prompting offers convenience, it underperformed in comparison to text based inputs, particularly in smaller models. Further improvements are expected with refined prompt engineering and few-shot examples. Refinement may lead to LLMs serving as the real time diagnostic assistants in esophageal motility evaluation. A key strength of our approach is the use of open-source models deployable locally, preserving patient privacy.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- S1034 Automated Interpretation of EndoFLIP Metrics Per Dallas Consensus Using Text- and Image-Based LLM Prompting
- Date Crossref
- 01/10/2025
- Éditeur
- Ovid Technologies (Wolters Kluwer Health)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.