An FPGA-Based Top-K Gradient Compression Accelerator for Distributed Deep Learning Training
Résumé fourni par la source
In distributed neural network training with multiple machines and devices, communication limitations often create efficiency bottlenecks due to the frequent exchange of model parameters and gradient information between computing nodes. This paper proposes an FPGA-based accelerator leveraging the gradient compression algorithm Top-K sparsification, which enhances performance by offloading computationally intensive compression operations to the FPGA. Experimental results demonstrate that the FPGA compression accelerator designed in this study achieves superior computing performance and compression efficiency compared to compression algorithms implemented on CPUs and GPUs. Specifically, the FPGA compresses the same amount of data 3.3-3.7 times faster than parallel solutions on CPUs and 1.3-1.8 times faster than on GPUs.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- An FPGA-Based Top-K Gradient Compression Accelerator for Distributed Deep Learning Training
- Date Crossref
- 22/10/2024
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.