Multi-Granularity Masking Strategy for Self-Supervised Visual Representation Learning
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Masked autoencoders have recently emerged as a powerful paradigm for self-supervised visual representation learning by reconstructing missing image patches. However, the widely adopted random masking mechanism ignores the inherent structural organization of visual data, resulting in limited interpretability and poor integration of global and local features. To overcome these limitations, we propose a multi-granularity masking strategy that hierarchically selects image patches from coarse to fine granularity. At higher levels, coarse-grained patches preserve global semantics, while finer-grained partitions progressively refine local details. This structured selection not only leverages semantic priors to guide the reconstruction process, but also ensures balanced coverage of contextual information across different scales. Experiments on the ImageNet-1K benchmark dataset demonstrate that our method achieves comparable or even superior downstream performance to the original MAE. Importantly, attention visualization analysis demonstrates that our strategy produces clearer, semantically coherent, and more focused attention maps, significantly enhancing the interpretability of the learned representations. This strategy highlights the potential of incorporating multi-granularity principles into self-supervised learning, opening up new avenues for interpretable and efficient visual transformers.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Multi-Granularity Masking Strategy for Self-Supervised Visual Representation Learning
- Date Crossref
- 21/11/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Chongqing University of Posts and Telecommunications pays non établi dans la noticeUniversité ou école supérieure
-
Chinese Academy of Sciences pays non établi dans la noticeOrganisme public
-
Shenzhen Institutes of Advanced Technology pays non établi dans la noticeStructure de recherche
-
Guangdong Police College pays non établi dans la noticeUniversité ou école supérieure
-
Sichuan Police College pays non établi dans la noticeUniversité ou école supérieure
Chongqing University of Posts and Telecommunications, Chinese Academy of Sciences et Shenzhen Institutes of Advanced Technology, avec 2 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.