LHA: Layer-wise Hardware Acceleration of Progressive Quantizing Inference through Partial Reconfiguration for Edge Computing

June 30, 2025·
Zongcheng Yue
Dr Sean Longyu Ma
Dr Sean Longyu Ma
,
Chiu-Wing Sham
,
Chong Fu
· 0 min read
Abstract
As the need for real-time, low-power deep learning at the edge increases, efficient hardware acceleration becomes crucial. Traditional edge hardware designs often scale to accommodate neural network sizes, which can degrade overall performance by taxing the hardware. To solve this, we propose a novel Layer-wise Hardware Acceleration (LHA) approach for Deep Neural Network (DNN) inference, leveraging progressive quantization and Partial Reconfiguration (PR). We first apply progressive quantization to systematically reduce the bit-width of network weights and activations, lowering computational and memory demands. Then, we utilize Field Programmable Gate Arrays (FPGAs) with PR capabilities to dynamically reconfigure hardware for each quantized network layer in sequence. This method optimizes FPGA resource usage, tailors to each layer’s needs, and reallocates freed resources to boost overall performance. Experiments show that LHA significantly enhances resource efficiency while maintaining inference performance on edge devices.
Type
Publication
2025 International Joint Conference on Neural Networks (IJCNN)
publications
Dr Sean Longyu Ma
Authors
Lecturer in Computer Science
Sean Longyu Ma is a Lecturer in the School of Computer Science at the University of Auckland. His research focuses on FPGA-based computing, RISC-V customisation, high-level synthesis, and heterogeneous computing.