Semi-SwinUNeTR: Towards 3D Swin Vision Transformer-Based UNet for Medical Image Segmentation with Limited Annotations
Rattachement africain : cn, gb. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for computer-assisted diagnosis, treatment planning, and disease monitoring. However, brain tumors usually exhibit irregular, heterogeneous, and multi-scale spatial patterns with complex and ambiguous boundaries. At the same time, the performance of deep segmentation models is often constrained by the limited availability of voxel-level annotations, which are expensive and time-consuming to obtain. To address these challenges, this paper proposes Semi-SwinUNeTR, a semi-supervised framework for 3D brain tumor segmentation with limited annotated data. The proposed method adopts SwinUNeTR as the segmentation backbone, enabling hierarchical volumetric representation learning through shifted-window self-attention while preserving the encoder–decoder structure required for dense prediction. On top of this backbone, we introduce a dual-consistency semi-supervised learning strategy, consisting of mean teacher-based model consistency and interpolation consistency-based data consistency. In addition, voxel-wise consistency weights are used to redistribute semi-supervised supervision toward structurally complex and boundary-irregular tumor regions without changing the SwinUNeTR backbone. Experiments on the BraTS 2019 benchmark demonstrate that the proposed framework achieves strong performance across different annotation ratios. The original Semi-SwinUNeTR achieves Dice scores of 84.93%, 86.25%, 87.05%, and 87.83% under the 10%, 20%, 40%, and 80% labeled-data settings, respectively. With the weighted consistency extension, the Dice scores are further improved to 85.64%, 87.94%, and 88.59% under the 10%, 20%, and 80% labeled-data settings, respectively, while the corresponding HD95 values are reduced to 8.9826, 8.1854, and 7.4533. These results indicate that combining a SwinUNeTR backbone with complementary model consistency, data consistency, and voxel-wise consistency weighting is an effective strategy for semi-supervised volumetric medical image segmentation under limited annotation.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Semi-SwinUNeTR: Towards 3D Swin Vision Transformer-Based UNet for Medical Image Segmentation with Limited Annotations
- Date Crossref
- 17/06/2026
- Éditeur
- MDPI AG
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Beijing University of Posts and Telecommunications Engineering Research Center of Blockchain and Network Convergence Technology pays non établi dans la noticeUniversité ou école supérieure
-
Aston University pays non établi dans la noticeUniversité ou école supérieure
-
School of Artificial Intelligence pays non établi dans la noticeUniversité ou école supérieure
-
School of Computer Science and Digital Technologies pays non établi dans la noticeUniversité ou école supérieure
Engineering Research Center of Blockchain and Network Convergence Technology — Beijing University of Posts and Telecommunications, Aston University et School of Artificial Intelligence, avec 1 autre affiliation.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.