Deploy Efficient Large Language Model Distributed Inference Pipeline for Heterogeneous GPUs
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
The advent of a large language model (LLM) has revolutionized various domains and services. The inference pipeline system is emerging as an efficient mechanism to deploy LLMs. However, existing works barely study the deployment of LLM inference on heterogeneous GPUs (with different computation and memory capabilities), where inference efficiency can be heavily affected by the imbalanced performance of different pipeline stages. Based on our empirical experience, the unbalanced pipeline stages incur GPU wait time, and average idle time can exceed 50% of the whole LLM inference process. In this paper, we study and optimize the distributed pipeline parallelism system for LLM inference on heterogeneous GPUs. We present a heuristic algorithm and implement a system that automatically deploys an efficient inference pipeline on heterogeneous GPUs. Extensive experiments are evaluated on 26 heterogeneous GPUs. The results demonstrate the superiority of our proposed system, which improves makespan (i.e., the total LLM inference latency) and throughput by an average of 37.1% and a maximum of 83.0% compared to the baselines.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Deploy Efficient Large Language Model Distributed Inference Pipeline for Heterogeneous GPUs
- Date Crossref
- 02/07/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
University of Science and Technology of China pays non établi dans la noticeUniversité ou école supérieure
University of Science and Technology of China.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.