Aller au contenu principal
2025 conference-paper

A Single-stage Interpretable Vision Transformer Model Via Granular-ball Computing

0Citations signalées, ce qui n’est pas une note de qualité
4Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Deep learning models have been widely used in image recognition tasks due to their superior feature learning capabilities. The Vision Transformer (ViT), a mainstream deep learning architecture, uses a self-attention mechanism that effectively captures global feature dependencies compared to traditional convolutional neural networks. However, existing ViT models, because of their single-scale nature, suffer from poor interpretability when target features vary in location and size. This paper proposes a purity-based multi-granularity single-stage transformer model designed to precisely locate key regions that influence classification decisions. This approach first uses a convolutional network to decompose the input image into multiple subimage patches of varying granularity, which are then fed into the ViT module for feature extraction and classification. Using the output of the attention weight matrix from the attention layer, this paper constructs an undirected graph structure with high-attention image patches as nodes. A single-stage random walk algorithm is then used to generate interpretable analysis results. Compared to traditional multi-stage random walk methods, this single-stage strategy significantly reduces computational complexity while effectively maintaining the advantages of multi-granularity feature importance analysis and enabling the quantification of the importance of feature combinations.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
A Single-stage Interpretable Vision Transformer Model Via Granular-ball Computing
Date Crossref
21/11/2025
Éditeur
IEEE
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Chongqing University of Posts and Telecommunications pays non établi dans la notice
    Université ou école supérieure
  • Chinese Academy of Sciences pays non établi dans la notice
    Organisme public
  • Shenzhen Institutes of Advanced Technology pays non établi dans la notice
    Structure de recherche
  • Guangdong Police College pays non établi dans la notice
    Université ou école supérieure
  • Sichuan Police College pays non établi dans la notice
    Université ou école supérieure

Chongqing University of Posts and Telecommunications, Chinese Academy of Sciences et Shenzhen Institutes of Advanced Technology, avec 2 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Ferroelectric and Negative Capacitance DevicesAdvanced Neural Network ApplicationsAdvanced Memory and Neural Computing

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.