PASS: A Priority-based Model Assignment for Minimal Inference Time in Serverless Edge Cloud
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Serverless computing is increasingly being adopted to provision various on-demand services at the edge cloud, including inference tasks based on deep neural networks (DNNs) for the Internet of Things (IoT). This approach leverages the advantages of flexible resource allocation and fine-grained resource management. However, the provisioning of on-demand inference typically requires downloading the DNN model at runtime, which can introduce significant delays. In the edge cloud with heterogeneous network connections, the inevitable model downloading time and the inter-model data transmission impose high challenges to the QoS of inference tasks. In this paper, we investigate how to jointly consider both model downloading time and communication time to minimize inference time. We first formulate this problem into a nonlinear optimization form and proved it as NP-hard. We further propose a Priority-based Model Assignment (PASS) algorithm in polynomial time and trace-driven experimental results show that it reduces the average inference time by 23.6% compared to existing state-of-the-art solutions.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- PASS: A Priority-based Model Assignment for Minimal Inference Time in Serverless Edge Cloud
- Date Crossref
- 04/08/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.