Using deep compression on PyTorch models for autonomous systems

Title Using deep compression on PyTorch models for autonomous systems
Author Doğan, Eren, Uğurdağ, Hasan Fatih, Ünlü, Hasan
Publication Date: 2022
Publication Place - IEEE
Type Document
Language English
Digital Yes
Manuscript No
Library: Özyeğin University
Library Asset ID 978-166545092-8
Record ID e84b9b68-0a27-4930-bec3-46f6fc5df6cc
Library Location Electrical & Electronics Engineering
Date 2022
Sample Text Applications of artificial neural networks on low-cost embedded systems and microcontrollers (MCUs), has recently been attracting more attention than ever. Since MCUs have limited memory capacity as well as limited compute-speed compared to workstations, employment of current deep learning algorithms on MCUs becomes more practical with the help of model compression. This makes MCUs common and practical alternative solution for autonomous systems. In this paper, we add model compression, specifically Deep Compression, to an existing work, which efficiently deploys PyTorch models on MCUs, in order to increase neural network speed and save electrical power. First, we prune the weight values close to zero in convolutional and fully connected layers. Secondly, the remaining weights and activations are quantized to 8-bit integers from 32-bit floating-point. Finally, forward pass functions are compressed using special data structures for sparse matrices, which store only nonzero weights. In the case of the LeNet-5 model, the memory footprint was reduced by 12.5x, and the inference speed was boosted by 2.6x.
DOI 10.1109/SIU55565.2022.9864848
View in source Özyeğin University Özyeğin University - Historical works, archives, and periodicals search engine
Özyeğin University - Historical works, archives, and periodicals search engine Özyeğin University

Using deep compression on PyTorch models for autonomous systems

Author Doğan, Eren, Uğurdağ, Hasan Fatih, Ünlü, Hasan
Publication Date 2022
Publication Place - IEEE
Type Document
Language English
Digital Yes
Manuscript No
Library Özyeğin University
Library Asset ID 978-166545092-8
Record ID e84b9b68-0a27-4930-bec3-46f6fc5df6cc
Library Location Electrical & Electronics Engineering
Date 2022
Sample Text Applications of artificial neural networks on low-cost embedded systems and microcontrollers (MCUs), has recently been attracting more attention than ever. Since MCUs have limited memory capacity as well as limited compute-speed compared to workstations, employment of current deep learning algorithms on MCUs becomes more practical with the help of model compression. This makes MCUs common and practical alternative solution for autonomous systems. In this paper, we add model compression, specifically Deep Compression, to an existing work, which efficiently deploys PyTorch models on MCUs, in order to increase neural network speed and save electrical power. First, we prune the weight values close to zero in convolutional and fully connected layers. Secondly, the remaining weights and activations are quantized to 8-bit integers from 32-bit floating-point. Finally, forward pass functions are compressed using special data structures for sparse matrices, which store only nonzero weights. In the case of the LeNet-5 model, the memory footprint was reduced by 12.5x, and the inference speed was boosted by 2.6x.
DOI 10.1109/SIU55565.2022.9864848
Özyeğin University - Historical works, archives, and periodicals search engine
Özyeğin University You are being redirected...

Please wait