TY - GEN
T1 - PriME
T2 - 43rd International Conference on Computer Design, ICCD 2025
AU - Lee, Junghyeok
AU - Jang, Jihoon
AU - Kim, Hyun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - With the growing demand for on-device AI, increasing efforts have been directed toward deploying lightweight small-scale large language models (sLLMs) on edge and mobile devices to enhance inference performance while minimizing computational cost and latency. As the number of decoder layers in sLLMs decreases, the embedding layer constitutes a substantial portion of the model's overall parameters and memory consumption. Consequently, efficient data compression is crucial; however, existing methods, such as quantization and pruning, drastically degrade accuracy when applied to the error-sensitive embedding layers. Moreover, embedding layer computations exhibit low arithmetic intensity (operations per byte), rendering them memory-bound. This limitation necessitates a shift from conventional von Neumann architectures to processing-in-memory (PIM) architectures. To address these challenges, this paper proposes (1) XOR-based Masking Compression (XMC), a lossless compression algorithm specialized for sLLM embedding layers, and (2) PriME, which integrates XMC with PIM architecture to alleviate memory bottlenecks in embedding layers. XMC enhances zero-bit representation in 16-bit FP data using ADD and XOR masking, achieving an average compression ratio of 1.49 × while being implementable with a 3-cycle decompression delay. PriME enables parallel processing of compressed data within PIM, accelerating embedding computations by an average of 4.0 × and up to 5.26 × compared to GPUs, while simultaneously reducing energy consumption by over 30 %, leading to an average energy efficiency improvement of 6.29 ×. Designed for broad applicability, PriME is compatible with various sLLMs and holds scalability for extension to multimodal small vision language models, demonstrating its versatility for efficient AI acceleration.
AB - With the growing demand for on-device AI, increasing efforts have been directed toward deploying lightweight small-scale large language models (sLLMs) on edge and mobile devices to enhance inference performance while minimizing computational cost and latency. As the number of decoder layers in sLLMs decreases, the embedding layer constitutes a substantial portion of the model's overall parameters and memory consumption. Consequently, efficient data compression is crucial; however, existing methods, such as quantization and pruning, drastically degrade accuracy when applied to the error-sensitive embedding layers. Moreover, embedding layer computations exhibit low arithmetic intensity (operations per byte), rendering them memory-bound. This limitation necessitates a shift from conventional von Neumann architectures to processing-in-memory (PIM) architectures. To address these challenges, this paper proposes (1) XOR-based Masking Compression (XMC), a lossless compression algorithm specialized for sLLM embedding layers, and (2) PriME, which integrates XMC with PIM architecture to alleviate memory bottlenecks in embedding layers. XMC enhances zero-bit representation in 16-bit FP data using ADD and XOR masking, achieving an average compression ratio of 1.49 × while being implementable with a 3-cycle decompression delay. PriME enables parallel processing of compressed data within PIM, accelerating embedding computations by an average of 4.0 × and up to 5.26 × compared to GPUs, while simultaneously reducing energy consumption by over 30 %, leading to an average energy efficiency improvement of 6.29 ×. Designed for broad applicability, PriME is compatible with various sLLMs and holds scalability for extension to multimodal small vision language models, demonstrating its versatility for efficient AI acceleration.
KW - embedding layer
KW - lossless compression
KW - on-device AI
KW - processing-in-memory
KW - quantization
KW - small LLMs
UR - https://www.scopus.com/pages/publications/105032523543
U2 - 10.1109/ICCD65941.2025.00038
DO - 10.1109/ICCD65941.2025.00038
M3 - Conference contribution
AN - SCOPUS:105032523543
T3 - Proceedings - IEEE International Conference on Computer Design: VLSI in Computers and Processors
SP - 221
EP - 228
BT - Proceedings - 2025 IEEE 43rd International Conference on Computer Design, ICCD 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 10 November 2025 through 12 November 2025
ER -