Abstract
Transformer-based deep neural networks (DNNs) have achieved remarkable success across a wide range of applications such as image and text generation tasks. However, The continuous growth of model size and memory demands imposes significant pressure on energy consumption and memory bandwidth, especially during weight access operations. To deal with such a challenge, prior studies have investigated NAND flash-based processing-in-memory (PIM) architectures, but it still experiences large energy consumption due to the significant increase in the recent model sizes. In this work, we present E-Flash, a digital NAND flash-based architecture for energy-efficient DNN weight access. E-Flash introduces a novel state-switching algorithm that reallocates frequently occurring weight patterns to low-power cell states in triple-level cell (TLC) flash memory. In addition, a cell-first allocation scheme further amplifies energy savings by aligning bit patterns within cells. Evaluation results on quantized BERT and Llama 2 models demonstrate up to 37.73% and 16.74% reduction in read energy, respectively, with negligible hardware overhead.
| Original language | English |
|---|---|
| Pages (from-to) | 1124-1133 |
| Number of pages | 10 |
| Journal | IEEE Transactions on Very Large Scale Integration (VLSI) Systems |
| Volume | 34 |
| Issue number | 4 |
| DOIs | |
| State | Published - 1 Apr 2026 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 7 Affordable and Clean Energy
Keywords
- NAND flash
- quadraple-level cell (QLC)
- state-switching algorithm
- transformers
- triple-level cell (TLC)
Fingerprint
Dive into the research topics of 'E-Flash: Energy-Efficient LLM Mapping on NAND Flash-Based In-Storage Inference Computing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver