TY - GEN
T1 - Memory-Efficient Depthwise Convolution Accelerator with Run-Length Coding
AU - Lee, Chaebin
AU - Kim, Jaeseong
AU - Lee, Dayoung
AU - Lee, Seung Eun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - In resource-constrained environments such as embedded systems, AI computations are limited by storage capacity and data transfer overhead. This paper proposes a hardware accelerator to reduce the memory usage of activation data for depthwise convolution in MobileNet-V1. In comparison with various compression algorithms, run-length coding (RLC) showed the highest compression ratio in almost all layers, and the memory usage was reduced by up to 69.29% when using RLC compared to uncompressed data. The hardware accelerator performs not only computation but also compression and decompression, whereas the software measures only computation. Nevertheless, the total execution time of the hardware was lower than that of the software. These results demonstrate that the proposed hardware accelerator can reduce memory usage without incurring processing time overhead from compression and decompression, making it suitable for embedded AI applications.
AB - In resource-constrained environments such as embedded systems, AI computations are limited by storage capacity and data transfer overhead. This paper proposes a hardware accelerator to reduce the memory usage of activation data for depthwise convolution in MobileNet-V1. In comparison with various compression algorithms, run-length coding (RLC) showed the highest compression ratio in almost all layers, and the memory usage was reduced by up to 69.29% when using RLC compared to uncompressed data. The hardware accelerator performs not only computation but also compression and decompression, whereas the software measures only computation. Nevertheless, the total execution time of the hardware was lower than that of the software. These results demonstrate that the proposed hardware accelerator can reduce memory usage without incurring processing time overhead from compression and decompression, making it suitable for embedded AI applications.
KW - depthwise convolution
KW - hardware accelerator
KW - RLC
UR - https://www.scopus.com/pages/publications/105033143876
U2 - 10.1109/ISOCC66390.2025.11329619
DO - 10.1109/ISOCC66390.2025.11329619
M3 - Conference contribution
AN - SCOPUS:105033143876
T3 - International SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers
BT - International SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 22nd International SoC Design Conference, ISOCC 2025
Y2 - 15 October 2025 through 18 October 2025
ER -