Aller au contenu principal
2025 conference-paper

Research on deep learning model reasoning optimization and parallel computing strategies under domesticated arithmetic architecture

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

With the continuous expansion of AI application scenarios, the demand of deep learning models for computing power resources continues to grow. In order to realize independent control of key technologies, localized arithmetic architecture has gradually become an important support in the field of artificial intelligence. However, limited by the differences in hardware instruction sets, compiler ecology and software compatibility, the inference efficiency of existing deep learning models on localized platforms is generally not high, facing performance bottlenecks and adaptation problems. In this paper, we focus on model inference optimization and parallel computing strategy under localized computing power architecture, aiming to improve the performance and deployment efficiency of deep learning tasks on local chips. Aiming at the hardware characteristics of localized chips, this paper first analyzes their key features in terms of computational architecture, memory structure, and supported operators, and identifies the core issues that constrain the inference performance. Subsequently, we propose targeted inference optimization strategies from multiple dimensions, such as operator fusion, memory access optimization, model compression, accuracy quantization, etc., and combine them with mainstream homegrown compiler platforms (e.g., MindSpore, Cambricon Neuware) for efficient runtime adaptation. Further, this paper designs a set of parallel computing and scheduling mechanisms suitable for domestic environments, covering data parallelism, model parallelism, and heterogeneous collaboration, which effectively improves the execution efficiency of multi-node distributed tasks.The experimental part of the performance test is based on the Rise and Cambrian platforms, and the results show that the optimization scheme proposed in this paper can significantly improve the reasoning speed and resource utilization while maintaining the model accuracy, demonstrating good engineering feasibility and application prospects. This research provides technical reference and realization path for building a high-performance and migratable domestic AI inference system.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Research on deep learning model reasoning optimization and parallel computing strategies under domesticated arithmetic architecture
Date Crossref
01/10/2025
Éditeur
Institution of Engineering and Technology (IET)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Educational Technology and Assessment

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.