EBPC: Extended Bit-Plane Compression for Deep Neural Network Inference\n and Training Accelerators
Le résumé fourni par la source
In the wake of the success of convolutional neural networks in image\nclassification, object recognition, speech recognition, etc., the demand for\ndeploying these compute-intensive ML models on embedded and mobile systems with\ntight power and energy constraints at low cost, as well as for boosting\nthroughput in data centers, is growing rapidly. This has sparked a surge of\nresearch into specialized hardware accelerators. Their performance is typically\nlimited by I/O bandwidth, power consumption is dominated by I/O transfers to\noff-chip memory, and on-chip memories occupy a large part of the silicon area.\nWe introduce and evaluate a novel, hardware-friendly, and lossless compression\nscheme for the feature maps present within convolutional neural networks. We\npresent hardware architectures and synthesis results for the compressor and\ndecompressor in 65nm. With a throughput of one 8-bit word/cycle at 600MHz, they\nfit into 2.8kGE and 3.0kGE of silicon area, respectively - together the size of\nless than seven 8-bit multiply-add units at the same throughput. We show that\nan average compression ratio of 5.1x for AlexNet, 4x for VGG-16, 2.4x for\nResNet-34 and 2.2x for MobileNetV2 can be achieved - a gain of 45-70% over\nexisting methods. Our approach also works effectively for various number\nformats, has a low frame-to-frame variance on the compression ratio, and\nachieves compression factors for gradient map compression during training that\nare even better than for inference.\n
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.