TY - GEN
T1 - Efficient MXINT4 Inference via Hardware-Based Dynamic Sparsity Detection and Skip-Activation
AU - Kim, Youngchan
AU - Kim, Hyun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Microscaling (MX) quantization has recently garnered significant attention due to its ability to achieve high compression ratios while preserving critical information across various deep learning tasks. It has been officially adopted and deployed across diverse hardware platforms. However, the blockwise processing required during MX quantization introduces considerable overhead, resulting in substantial latency when executed on conventional hardware. Consequently, it becomes challenging to realize actual computational acceleration using standard architectures. This highlights the necessity for a dedicated hardware design tailored to MX quantization, capable of mitigating such overhead and enabling practical speedups. In this paper, we present a hardware quantizer module specifically designed for the MX format and propose the MX quantizer with dynamic sparsity detection (MQDS), which identifies activation sparsity during the quantization process. Furthermore, we introduce a sparse computation flow that leverages the outputs of MQDS to address the area overhead and latency compatibility issues inherent in conventional MX quantization approaches. Experimental results demonstrate that MQDS achieves an average activation skip detection ratio of 66.49 % on vision transformer models while occupying only 12.2 % of the area compared to traditional integer quantization designs.
AB - Microscaling (MX) quantization has recently garnered significant attention due to its ability to achieve high compression ratios while preserving critical information across various deep learning tasks. It has been officially adopted and deployed across diverse hardware platforms. However, the blockwise processing required during MX quantization introduces considerable overhead, resulting in substantial latency when executed on conventional hardware. Consequently, it becomes challenging to realize actual computational acceleration using standard architectures. This highlights the necessity for a dedicated hardware design tailored to MX quantization, capable of mitigating such overhead and enabling practical speedups. In this paper, we present a hardware quantizer module specifically designed for the MX format and propose the MX quantizer with dynamic sparsity detection (MQDS), which identifies activation sparsity during the quantization process. Furthermore, we introduce a sparse computation flow that leverages the outputs of MQDS to address the area overhead and latency compatibility issues inherent in conventional MX quantization approaches. Experimental results demonstrate that MQDS achieves an average activation skip detection ratio of 66.49 % on vision transformer models while occupying only 12.2 % of the area compared to traditional integer quantization designs.
KW - Dynamic sparsity
KW - Microscaling (MX)
KW - MX quantization
KW - Sparse computation
UR - https://www.scopus.com/pages/publications/105035387757
U2 - 10.1109/APCCAS67402.2025.11377131
DO - 10.1109/APCCAS67402.2025.11377131
M3 - Conference contribution
AN - SCOPUS:105035387757
T3 - Proceedings - 2025 21st IEEE Asia Pacific Conference on Circuits and Systems, APCCAS 2025
BT - Proceedings - 2025 21st IEEE Asia Pacific Conference on Circuits and Systems, APCCAS 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 21st IEEE Asia Pacific Conference on Circuits and Systems, APCCAS 2025
Y2 - 12 October 2025 through 15 October 2025
ER -