Research on deep learning model reasoning optimization and parallel computing strategies under domesticated arithmetic architecture
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
With the continuous expansion of AI application scenarios, the demand of deep learning models for computing power resources continues to grow. In order to realize independent control of key technologies, localized arithmetic architecture has gradually become an important support in the field of artificial intelligence. However, limited by the differences in hardware instruction sets, compiler ecology and software compatibility, the inference efficiency of existing deep learning models on localized platforms is generally not high, facing performance bottlenecks and adaptation problems. In this paper, we focus on model inference optimization and parallel computing strategy under localized computing power architecture, aiming to improve the performance and deployment efficiency of deep learning tasks on local chips. Aiming at the hardware characteristics of localized chips, this paper first analyzes their key features in terms of computational architecture, memory structure, and supported operators, and identifies the core issues that constrain the inference performance. Subsequently, we propose targeted inference optimization strategies from multiple dimensions, such as operator fusion, memory access optimization, model compression, accuracy quantization, etc., and combine them with mainstream homegrown compiler platforms (e.g., MindSpore, Cambricon Neuware) for efficient runtime adaptation. Further, this paper designs a set of parallel computing and scheduling mechanisms suitable for domestic environments, covering data parallelism, model parallelism, and heterogeneous collaboration, which effectively improves the execution efficiency of multi-node distributed tasks.The experimental part of the performance test is based on the Rise and Cambrian platforms, and the results show that the optimization scheme proposed in this paper can significantly improve the reasoning speed and resource utilization while maintaining the model accuracy, demonstrating good engineering feasibility and application prospects. This research provides technical reference and realization path for building a high-performance and migratable domestic AI inference system.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Research on deep learning model reasoning optimization and parallel computing strategies under domesticated arithmetic architecture
- Date Crossref
- 01/10/2025
- Éditeur
- Institution of Engineering and Technology (IET)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.