Aller au contenu principal
2025 article

Variation-Aware Demonstration and Optimization of Block Floating Point and Integer Neural Network Acceleration on RRAM Compute in-Memory Hardware

2Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : in, tw. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

In this work, we propose a device-circuit-system co-design approach to achieve the best performance from a given variation-prone resistive random access memory (RRAM) compute-in-memory (CIM) hardware for deep neural network (DNN) acceleration and also demonstrate an optimized hardware implementation of block floating point (BFP) acceleration exploiting FP8 format on RRAM CIM hardware using a 1T-1R array of single-level cell (SLC) and multilevel cell (MLC) devices. Further, bit-level accurate network simulations incorporating device-to-device (D2D) variations on multi-bit SLC / MLC weight mantissa / weight are also conducted and benchmarking is performed with integer (INT4, INT8) formats. In particular, a word line read voltage ($V_{\boldsymbol{WLread}}$) tuning scheme is proposed and experimentally analyzed to (a) optimize D2D variation (tested on 144 1T-1R array bit cells) in SLC-based implementation to reduce partial MAC (PMAC) mapping errors for robust inference and b) obtain linearly separated conductance states for linear accumulation in MLC-based implementation (tested on 200 1T-1R bitcells). An optimized write-verify scheme is also proposed to obtain a tight MLC states distribution. Extensive analysis on hardware obtained FP8 inference results is also performed by measuring 16,384$<$PMACs$<$1,31,072 from hardware to identify sources of error for different word-line parallelism, SLC / MLC mapping and their trade-off with sparsity and energy consumption. For SLC based hardware implementation, the test accuracy of the FP8 trained networks at$\sigma/\mu =$0.5 is 1.2x better than INT4 at only${\boldsymbol\sim}$1.3x the area consumption of the INT4 RRAM array. INT8 network performance on the other hand, is comparable to FP8 but at${\boldsymbol\sim}$2x FP8 RRAM array area consumption. Hardware inference using 2-bit MLC mapped weight mantissas is also demonstrated to analyze the effect of individual state variation dependence and identify critical variability states.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Variation-Aware Demonstration and Optimization of Block Floating Point and Integer Neural Network Acceleration on RRAM Compute in-Memory Hardware
Date Crossref
01/09/2025
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Indian Institute of Technology Delhi Department of Electrical Engineering pays non établi dans la notice
    Université ou école supérieure
  • National Yang Ming Chiao Tung University Department of Electrical Engineering and Computer Science pays non établi dans la notice
    Université ou école supérieure

Department of Electrical Engineering — Indian Institute of Technology Delhi et Department of Electrical Engineering and Computer Science — National Yang Ming Chiao Tung University.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Advanced Memory and Neural ComputingFerroelectric and Negative Capacitance DevicesAdvanced Neural Network Applications

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.