TY - GEN
T1 - RANQ
T2 - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
AU - Jeong, Sangbeom
AU - Kim, Hyun
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Large language models (LLMs) have demonstrated remarkable performance across a wide range of domains. However, their deployment in real-world environments remains challenging due to the substantial model sizes and limited hardware resources. Quantization has been widely adopted to alleviate these issues, but the presence of outliers in the weight distribution of LLMs significantly degrades quantization performance, particularly in low-bit settings. To address this challenge, we propose a region-adaptive non-uniform quantization (RANQ) method that takes into account the intrinsic characteristics of LLM weight distributions. RANQ divides the entire weight space into central and outlier regions, applying distinct scale factors to each region to minimize performance degradation caused by quantization. Furthermore, we introduce a layer-wise reconstruction-based threshold optimization algorithm, which adaptively determines the optimal region boundaries according to the distributional properties of each layer, enabling more precise low-bit quantization. Experimental results on the LLaMA family using the WikiText2 and C4 datasets demonstrate that the proposed method consistently outperforms existing post training quantization (PTQ) techniques, such as round-to-nearest (RTN) and GPTQ.
AB - Large language models (LLMs) have demonstrated remarkable performance across a wide range of domains. However, their deployment in real-world environments remains challenging due to the substantial model sizes and limited hardware resources. Quantization has been widely adopted to alleviate these issues, but the presence of outliers in the weight distribution of LLMs significantly degrades quantization performance, particularly in low-bit settings. To address this challenge, we propose a region-adaptive non-uniform quantization (RANQ) method that takes into account the intrinsic characteristics of LLM weight distributions. RANQ divides the entire weight space into central and outlier regions, applying distinct scale factors to each region to minimize performance degradation caused by quantization. Furthermore, we introduce a layer-wise reconstruction-based threshold optimization algorithm, which adaptively determines the optimal region boundaries according to the distributional properties of each layer, enabling more precise low-bit quantization. Experimental results on the LLaMA family using the WikiText2 and C4 datasets demonstrate that the proposed method consistently outperforms existing post training quantization (PTQ) techniques, such as round-to-nearest (RTN) and GPTQ.
KW - Large language models
KW - low-bit quantization
KW - model compression
UR - https://www.scopus.com/pages/publications/105034892909
U2 - 10.1109/ICEIC69189.2026.11386208
DO - 10.1109/ICEIC69189.2026.11386208
M3 - Conference contribution
AN - SCOPUS:105034892909
T3 - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
BT - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 18 January 2026 through 21 January 2026
ER -