Abstract 2438: Bridging whole slide images and large language model for slide-level question answering
Résumé fourni par la source
Abstract Background: Vision Language Models (VLMs), integrating large language models (LLMs) with vision encoders to enable better understanding and interaction with visual data, have broken entirely new ground in computational pathology. Recent studies have demonstrated that employing a simple linear layer as an 'adaptor' to transform image embeddings, extracted by pathology-specific vision encoders, into LLM-recognizable tokens is effective for developing patch-level VLMs in pathology. However, the gigapixel-scale size of whole-slide images (WSIs) and their pyramidal structure pose significant challenges in designing effective adaptors for slide-level analysis. In particular, different diagnostic questions and tasks at the slide level require focusing on various sub-regions at different magnifications or combining multiple regions and magnifications. Here, we aim to assess adaptors that meet these challenging demands and develop a slide-level VLM to enable slide-level question answering. Methods: We develop Llama-Path, the first general-purpose slide-level VLM, using WSIs paired with descriptions and question-answer (Q&A) pairs. First, we collect 35,943 WSIs from public databases such as TCGA and GTEx, along with corresponding pathological reports or notes, covering a wide range of tumor and tissue types. Based on these reports and notes, we curate a WSI description dataset and a WSI Q&A dataset (337,000 Q&A pairs) using GPT-4. Then, we extract hierarchical patch embeddings from each WSI by the pathology foundation model, Conch. Next, we propose and evaluate six distinct slide-level adaptors: four vision-only, and two vision-text interactive. These adaptors are finetuned with Llama-3.1-8B across two stages to develop various versions of Llama-Path: The first stage uses the WSI description dataset to provide Llama-Path with a slide-level vision understanding ability. The second stage enhances the slide-level instruction-following capability of Llama-Path with the WSI Q&A dataset. Finally, we evaluate different versions of Llama-Path on a hold-out dataset consisting of 367 tumor and non-tumor cases with 4,440 Q&A pairs. Results: All versions of Llama-Path demonstrated exceptional performance, each achieving over 90% accuracy in close-ended Q&A pairs. GPT-4 assessed open-ended performances with high accuracy, showing 36% of responses correct, 27% partially correct, and 3% incorrect. The remaining 34% could not be evaluated without access to the corresponding WSIs. The LongNet-based vision-text interactive adaptor outperformed others, likely due to its targeted capture of visual information and simultaneous optimisation of WSI embeddings. Conclusions: Our findings highlight the potential of VLMs to enhance slide-level computer-aided diagnosis systems. Results from ongoing analyses, based on pathologists’ evaluations, will be presented shortly. Citation Format: Zeyu Gao, Kai He, Weiheng Su, William McGough, Ines Prata Machado, Mercedes Jimenez-Linan, Brian Rous, Rebecca Wray, Chengzu Li, Xiaobo Pang, Tieliang Gong, Mengling Feng, Chen Li, Mireia Crispin-Ortuzar. Bridging whole slide images and large language model for slide-level question answering [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2438.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Abstract 2438: Bridging whole slide images and large language model for slide-level question answering
- Date Crossref
- 21/04/2025
- Éditeur
- American Association for Cancer Research (AACR)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.