LHA: Layer-wise Hardware Acceleration of Progressive Quantizing Inference through Partial Reconfiguration for Edge Computing
June 30, 2025·
,,·
0 min read
Zongcheng Yue
Dr Sean Longyu Ma
Chiu-Wing Sham
Chong Fu
Abstract
As the need for real-time, low-power deep learning at the edge increases, efficient hardware acceleration becomes crucial. Traditional edge hardware designs often scale to accommodate neural network sizes, which can degrade overall performance by taxing the hardware. To solve this, we propose a novel Layer-wise Hardware Acceleration (LHA) approach for Deep Neural Network (DNN) inference, leveraging progressive quantization and Partial Reconfiguration (PR). We first apply progressive quantization to systematically reduce the bit-width of network weights and activations, lowering computational and memory demands. Then, we utilize Field Programmable Gate Arrays (FPGAs) with PR capabilities to dynamically reconfigure hardware for each quantized network layer in sequence. This method optimizes FPGA resource usage, tailors to each layer’s needs, and reallocates freed resources to boost overall performance. Experiments show that LHA significantly enhances resource efficiency while maintaining inference performance on edge devices.
Type
Publication
2025 International Joint Conference on Neural Networks (IJCNN)

Authors
Dr Sean Longyu Ma
(he/him)
Lecturer in Computer Science
Sean Longyu Ma is a Lecturer in the School of Computer Science at the
University of Auckland. His research focuses on FPGA-based computing,
RISC-V customisation, high-level synthesis, and heterogeneous computing.