TY - GEN
T1 - Toward Efficient Deployment of Mixture of Experts Models
T2 - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
AU - Lee, Dongjun
AU - Shin, Jin
AU - Kim, Hyun
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Recent mixture of experts (MoE)-based transformer models have achieved remarkable performance across various applications. However, as model size increases, the associated memory demand escalates dramatically, limiting their deployment in resource-constrained edge environments. To address this challenge, this study investigates the trade-off between performance and computational overhead under extreme compression conditions by applying ternary quantization and additional compression schemes to MoE models. Experimental results demonstrate that Huffman coding achieves a higher compression ratio but suffers from substantial decoding latency, while dictionary coding provides a lower compression ratio with significantly faster decoding. These findings highlight that efficient deployment of MoE models requires an integrated compression framework that jointly optimizes compression ratio and overhead. Furthermore, based on these observations, potential research directions are suggested, including hybrid coding formats and look-up table-based dedicated hardware accelerators for efficient MoE model compression.
AB - Recent mixture of experts (MoE)-based transformer models have achieved remarkable performance across various applications. However, as model size increases, the associated memory demand escalates dramatically, limiting their deployment in resource-constrained edge environments. To address this challenge, this study investigates the trade-off between performance and computational overhead under extreme compression conditions by applying ternary quantization and additional compression schemes to MoE models. Experimental results demonstrate that Huffman coding achieves a higher compression ratio but suffers from substantial decoding latency, while dictionary coding provides a lower compression ratio with significantly faster decoding. These findings highlight that efficient deployment of MoE models requires an integrated compression framework that jointly optimizes compression ratio and overhead. Furthermore, based on these observations, potential research directions are suggested, including hybrid coding formats and look-up table-based dedicated hardware accelerators for efficient MoE model compression.
KW - Dedicated Hardware Accelerators
KW - Edge Deploy
KW - Faster Decoding
KW - Mixture of Experts
KW - Ternary Quantization
UR - https://www.scopus.com/pages/publications/105034870532
U2 - 10.1109/ICEIC69189.2026.11386441
DO - 10.1109/ICEIC69189.2026.11386441
M3 - Conference contribution
AN - SCOPUS:105034870532
T3 - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
BT - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 18 January 2026 through 21 January 2026
ER -