A hybrid ConvNeXt–BiLSTM framework for robust scene text recognition
Rattachement africain : Égypte. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Scene Text Recognition (STR) is a fundamental computer vision task with broad applications in autonomous navigation, document digitization, and assistive technologies. However, traditional STR models often rely heavily on large synthetic datasets due to the scarcity of annotated real-world data, which limits their generalization in complex environments. To address this challenge, this study proposes a ConvNeXt-based deep learning framework that integrates Convolutional Network Next (ConvNeXt) for robust feature extraction with Bidirectional Long Short-Term Memory (BiLSTM) networks for effective sequence modeling. The framework incorporates label smoothing and focal loss to enhance training stability and alleviate class imbalance and overconfidence issues. Training is conducted in two stages: pre-training on synthetic datasets (MJSynth and SynthText) followed by fine-tuning on diverse real-world datasets, including IC13, IC15, RCTW, ArT, LSVT, MLT19, ReCTS, COCO-Text, Uber-Text, TextOCR, OpenVINO, and a subset of Union14M-L. Experimental results demonstrate that the proposed model achieves an average accuracy of 94.71% over six standard STR benchmarks (IIIT5k, SVT, IC13, IC15, SVTP, and CUTE80) when trained on both synthetic and real datasets, surpassing the 89.1% accuracy achieved using synthetic data alone on the same benchmarks, and outperforming state-of-the-art methods trained under comparable data conditions. The integration of ConvNeXt, BiLSTM, advanced loss functions, and heterogeneous datasets substantially improve STR performance, particularly under challenging conditions involving irregular text layouts, multilingual content, and complex backgrounds. Furthermore, the complete recognition pipeline achieves 20.3 M parameters, 1.9 GFLOPs, and an inference latency of 2.638 ms per image, demonstrating the practical suitability of the proposed framework for real-time deployment.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A hybrid ConvNeXt–BiLSTM framework for robust scene text recognition
- Date Crossref
- 13/05/2026
- Éditeur
- Springer Science and Business Media LLC
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.