YOLOv10-CBRC: A high-precision document image layout analysis model
Rattachement africain : bd, cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Tibetan document layout analysis (DLA) is a crucial aspect of the digitization of Tibetan documents. However, two significant issues persist: (1) the inadequacy of Tibetan document datasets;(2) insufficient utilization of layout information by existing models. To address these challenges, we have constructed a novel dataset, the Aba Tibetan Newspaper Dataset (AbaTND), consisting of 571 color images of Tibetan newspapers, effectively filling a data gap in this field. Additionally, we propose an advanced model, YOLOv10-CBRC (YOLOv10-CBAM-Re-Upsample-Re-SCDown-CRCV3), along with its variant YOLOv10-RC (YOLOv10-Re-Upsample-Re-SCDown-CRCV3), aimed at enhancing information utilization. Building upon YOLOv10 as the baseline model, the proposed model implements three key improvements: (1) the replacement of the Partial Self-Attention (PSA) module with the Convolutional Block Attention Module (CBAM), which enhances perception of channel-wise information and spatial localization of objects,(2) the adoption of dual-branch Re-upsample and Re-SCDown modules (2Re), which facilitates more effective utilization of multi-scale information,(3) the design of a novel classification feature processor, CIBwithResidualCV3 (CRCV3), which improves performance in classification tasks.Experimental results demonstrate that YOLOv10-CBRC achieves a mAP $$_{50\text {-}95}$$ of 77.6% on the AbaTND, while YOLOv10-RC reaches mAP $$_{50\text {-}95}$$ scores of 70.6% and 74.6% on the $$D^4LA$$ and IIIT-AR-13K datasets, respectively, significantly outperforming baseline model. The dataset and source code are publicly available at: https://github.com/fengmuyanghua/YOLOv10-CBRC-AbaTND .
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- YOLOv10-CBRC: A high-precision document image layout analysis model
- Date Crossref
- 31/07/2025
- Éditeur
- Springer Science and Business Media LLC
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.