Skip to main navigation Skip to search Skip to main content

Toward Efficient Deployment of Mixture of Experts Models: Quantization and Compression Analysis

  • Seoul National University of Science and Technology (SNUST)

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent mixture of experts (MoE)-based transformer models have achieved remarkable performance across various applications. However, as model size increases, the associated memory demand escalates dramatically, limiting their deployment in resource-constrained edge environments. To address this challenge, this study investigates the trade-off between performance and computational overhead under extreme compression conditions by applying ternary quantization and additional compression schemes to MoE models. Experimental results demonstrate that Huffman coding achieves a higher compression ratio but suffers from substantial decoding latency, while dictionary coding provides a lower compression ratio with significantly faster decoding. These findings highlight that efficient deployment of MoE models requires an integrated compression framework that jointly optimizes compression ratio and overhead. Furthermore, based on these observations, potential research directions are suggested, including hybrid coding formats and look-up table-based dedicated hardware accelerators for efficient MoE model compression.

Original languageEnglish
Title of host publication2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331580773
DOIs
StatePublished - 2026
Event2026 International Conference on Electronics, Information, and Communication, ICEIC 2026 - Macau, China
Duration: 18 Jan 202621 Jan 2026

Publication series

Name2026 International Conference on Electronics, Information, and Communication, ICEIC 2026

Conference

Conference2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
Country/TerritoryChina
CityMacau
Period18/01/2621/01/26

Keywords

  • Dedicated Hardware Accelerators
  • Edge Deploy
  • Faster Decoding
  • Mixture of Experts
  • Ternary Quantization

Fingerprint

Dive into the research topics of 'Toward Efficient Deployment of Mixture of Experts Models: Quantization and Compression Analysis'. Together they form a unique fingerprint.

Cite this