A 40 TOPS Single-Chip Accelerator Enabling Low-Latency Inference for Deep Neural Networks
Résumé fourni par la source
To achieve low latency for edge applications, a singlechip sparse accelerator is proposed, which can conduct deep neural network (DNN) inference only using limited on-chip memory. Private memory is eliminated, and all memories are shared to reduce power and chip area. An adaptive and variablelength compression algorithm is proposed to store sparse DNNs. A weak-constrained pruning algorithm is proposed to resolve load balance issue in kernel level, which can achieve almost the same sparsity as unconstrained pruning schemes (UCP). Based on these works, a low latency inference accelerator is fabricated in 28-nm CMOS with 8256 MACs and 9.4 MB on-chip SRAM, which can achieve a latency of 0.44 ms for YOLO3 tiny. For highsparsity layers, our chip can achieve 6.1× speedup and a throughput of 40 TOPS. With a pruned YOLO model, our accelerator achieves 6.7× lower latency and 21.7× better energy efficiency than Jetson Orin. A high-speed evaluation platform is built to demonstrate real-time object detection at a throughput of 600 frames per second (fps) with a power of 1.34 W.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A 40 TOPS Single-Chip Accelerator Enabling Low-Latency Inference for Deep Neural Networks
- Date Crossref
- 01/06/2025
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.