A Single-stage Interpretable Vision Transformer Model Via Granular-ball Computing
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Deep learning models have been widely used in image recognition tasks due to their superior feature learning capabilities. The Vision Transformer (ViT), a mainstream deep learning architecture, uses a self-attention mechanism that effectively captures global feature dependencies compared to traditional convolutional neural networks. However, existing ViT models, because of their single-scale nature, suffer from poor interpretability when target features vary in location and size. This paper proposes a purity-based multi-granularity single-stage transformer model designed to precisely locate key regions that influence classification decisions. This approach first uses a convolutional network to decompose the input image into multiple subimage patches of varying granularity, which are then fed into the ViT module for feature extraction and classification. Using the output of the attention weight matrix from the attention layer, this paper constructs an undirected graph structure with high-attention image patches as nodes. A single-stage random walk algorithm is then used to generate interpretable analysis results. Compared to traditional multi-stage random walk methods, this single-stage strategy significantly reduces computational complexity while effectively maintaining the advantages of multi-granularity feature importance analysis and enabling the quantification of the importance of feature combinations.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A Single-stage Interpretable Vision Transformer Model Via Granular-ball Computing
- Date Crossref
- 21/11/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Chongqing University of Posts and Telecommunications pays non établi dans la noticeUniversité ou école supérieure
-
Chinese Academy of Sciences pays non établi dans la noticeOrganisme public
-
Shenzhen Institutes of Advanced Technology pays non établi dans la noticeStructure de recherche
-
Guangdong Police College pays non établi dans la noticeUniversité ou école supérieure
-
Sichuan Police College pays non établi dans la noticeUniversité ou école supérieure
Chongqing University of Posts and Telecommunications, Chinese Academy of Sciences et Shenzhen Institutes of Advanced Technology, avec 2 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.