Variation-Aware Demonstration and Optimization of Block Floating Point and Integer Neural Network Acceleration on RRAM Compute in-Memory Hardware
Rattachement africain : in, tw. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
In this work, we propose a device-circuit-system co-design approach to achieve the best performance from a given variation-prone resistive random access memory (RRAM) compute-in-memory (CIM) hardware for deep neural network (DNN) acceleration and also demonstrate an optimized hardware implementation of block floating point (BFP) acceleration exploiting FP8 format on RRAM CIM hardware using a 1T-1R array of single-level cell (SLC) and multilevel cell (MLC) devices. Further, bit-level accurate network simulations incorporating device-to-device (D2D) variations on multi-bit SLC / MLC weight mantissa / weight are also conducted and benchmarking is performed with integer (INT4, INT8) formats. In particular, a word line read voltage ($V_{\boldsymbol{WLread}}$) tuning scheme is proposed and experimentally analyzed to (a) optimize D2D variation (tested on 144 1T-1R array bit cells) in SLC-based implementation to reduce partial MAC (PMAC) mapping errors for robust inference and b) obtain linearly separated conductance states for linear accumulation in MLC-based implementation (tested on 200 1T-1R bitcells). An optimized write-verify scheme is also proposed to obtain a tight MLC states distribution. Extensive analysis on hardware obtained FP8 inference results is also performed by measuring 16,384$<$PMACs$<$1,31,072 from hardware to identify sources of error for different word-line parallelism, SLC / MLC mapping and their trade-off with sparsity and energy consumption. For SLC based hardware implementation, the test accuracy of the FP8 trained networks at$\sigma/\mu =$0.5 is 1.2x better than INT4 at only${\boldsymbol\sim}$1.3x the area consumption of the INT4 RRAM array. INT8 network performance on the other hand, is comparable to FP8 but at${\boldsymbol\sim}$2x FP8 RRAM array area consumption. Hardware inference using 2-bit MLC mapped weight mantissas is also demonstrated to analyze the effect of individual state variation dependence and identify critical variability states.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Variation-Aware Demonstration and Optimization of Block Floating Point and Integer Neural Network Acceleration on RRAM Compute in-Memory Hardware
- Date Crossref
- 01/09/2025
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Indian Institute of Technology Delhi Department of Electrical Engineering pays non établi dans la noticeUniversité ou école supérieure
-
National Yang Ming Chiao Tung University Department of Electrical Engineering and Computer Science pays non établi dans la noticeUniversité ou école supérieure
Department of Electrical Engineering — Indian Institute of Technology Delhi et Department of Electrical Engineering and Computer Science — National Yang Ming Chiao Tung University.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.