2-D GAF-Enhanced Multimodal Vision Language Model for Breathing Patterns Analysis via ECG Sensing
Résumé fourni par la source
The use of vision large language models (VLLMs) for the analysis and diagnosis of respiratory patterns is affected by limited specialist training data and the lack of integrated medical expertise. This paper introduces a method that uses a multimodal VLLM to analyze electrocardiogram (ECG) data transformed to Gramian Angluar Fields (GAF) 2D image to determine respiratory patterns. Using advanced Shimmer3 sensing. Five VLLMs are fine-tuned using four different ECG datasets for training, each containing 5000 samples, to improve the interpretation of ECG signals. VLLMs are then evaluated against an independent set of data to determine normal, rapid, and difficult breathing patterns. The low-rank adaptation method is used to optimize VLLMs for the analysis of ECG signals obtained from the shimmer sensing platform. Results show that Idefics outperforms other fine-tuned VLLMs with a BLEU score of 0.572 for 2D ECG signals and 0.549 for 1D ECG signals, which demonstrates that the 2D image representation of the GAF improves the visualization of ECG respiratory patterns.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- 2-D GAF-Enhanced Multimodal Vision Language Model for Breathing Patterns Analysis via ECG Sensing
- Date Crossref
- 01/12/2025
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.