Aller au contenu principal
Accès ouvert déclaré 2025 article

YOLOv10-CBRC: A high-precision document image layout analysis model

3Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : bd, cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Tibetan document layout analysis (DLA) is a crucial aspect of the digitization of Tibetan documents. However, two significant issues persist: (1) the inadequacy of Tibetan document datasets;(2) insufficient utilization of layout information by existing models. To address these challenges, we have constructed a novel dataset, the Aba Tibetan Newspaper Dataset (AbaTND), consisting of 571 color images of Tibetan newspapers, effectively filling a data gap in this field. Additionally, we propose an advanced model, YOLOv10-CBRC (YOLOv10-CBAM-Re-Upsample-Re-SCDown-CRCV3), along with its variant YOLOv10-RC (YOLOv10-Re-Upsample-Re-SCDown-CRCV3), aimed at enhancing information utilization. Building upon YOLOv10 as the baseline model, the proposed model implements three key improvements: (1) the replacement of the Partial Self-Attention (PSA) module with the Convolutional Block Attention Module (CBAM), which enhances perception of channel-wise information and spatial localization of objects,(2) the adoption of dual-branch Re-upsample and Re-SCDown modules (2Re), which facilitates more effective utilization of multi-scale information,(3) the design of a novel classification feature processor, CIBwithResidualCV3 (CRCV3), which improves performance in classification tasks.Experimental results demonstrate that YOLOv10-CBRC achieves a mAP $$_{50\text {-}95}$$ of 77.6% on the AbaTND, while YOLOv10-RC reaches mAP $$_{50\text {-}95}$$ scores of 70.6% and 74.6% on the $$D^4LA$$ and IIIT-AR-13K datasets, respectively, significantly outperforming baseline model. The dataset and source code are publicly available at: https://github.com/fengmuyanghua/YOLOv10-CBRC-AbaTND .

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
YOLOv10-CBRC: A high-precision document image layout analysis model
Date Crossref
31/07/2025
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Handwritten Text Recognition TechniquesAdvanced Image and Video Retrieval TechniquesImage Retrieval and Classification Techniques

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.