Skip to main navigation Skip to search Skip to main content

PriME: PIM-Aware Efficient Compression for Memory-Bound Embedding Layers in sLLMs

  • Seoul National University of Science and Technology (SNUST)

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

With the growing demand for on-device AI, increasing efforts have been directed toward deploying lightweight small-scale large language models (sLLMs) on edge and mobile devices to enhance inference performance while minimizing computational cost and latency. As the number of decoder layers in sLLMs decreases, the embedding layer constitutes a substantial portion of the model's overall parameters and memory consumption. Consequently, efficient data compression is crucial; however, existing methods, such as quantization and pruning, drastically degrade accuracy when applied to the error-sensitive embedding layers. Moreover, embedding layer computations exhibit low arithmetic intensity (operations per byte), rendering them memory-bound. This limitation necessitates a shift from conventional von Neumann architectures to processing-in-memory (PIM) architectures. To address these challenges, this paper proposes (1) XOR-based Masking Compression (XMC), a lossless compression algorithm specialized for sLLM embedding layers, and (2) PriME, which integrates XMC with PIM architecture to alleviate memory bottlenecks in embedding layers. XMC enhances zero-bit representation in 16-bit FP data using ADD and XOR masking, achieving an average compression ratio of 1.49 × while being implementable with a 3-cycle decompression delay. PriME enables parallel processing of compressed data within PIM, accelerating embedding computations by an average of 4.0 × and up to 5.26 × compared to GPUs, while simultaneously reducing energy consumption by over 30 %, leading to an average energy efficiency improvement of 6.29 ×. Designed for broad applicability, PriME is compatible with various sLLMs and holds scalability for extension to multimodal small vision language models, demonstrating its versatility for efficient AI acceleration.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE 43rd International Conference on Computer Design, ICCD 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages221-228
Number of pages8
ISBN (Electronic)9798331503468
DOIs
StatePublished - 2025
Event43rd International Conference on Computer Design, ICCD 2025 - Richardson, United States
Duration: 10 Nov 202512 Nov 2025

Publication series

NameProceedings - IEEE International Conference on Computer Design: VLSI in Computers and Processors
ISSN (Print)1063-6404

Conference

Conference43rd International Conference on Computer Design, ICCD 2025
Country/TerritoryUnited States
CityRichardson
Period10/11/2512/11/25

Keywords

  • embedding layer
  • lossless compression
  • on-device AI
  • processing-in-memory
  • quantization
  • small LLMs

Fingerprint

Dive into the research topics of 'PriME: PIM-Aware Efficient Compression for Memory-Bound Embedding Layers in sLLMs'. Together they form a unique fingerprint.

Cite this