Skip to main navigation Skip to search Skip to main content

Towards efficient language giants: A comprehensive survey on structural optimizations and compression techniques for large language models

  • Gilhyeon Lee
  • , Seonggeun Kim
  • , Dongjun Lee
  • , Kyungmin Goh
  • , Hyun Kim
  • Seoul National University of Science and Technology (SNUST)

Research output: Contribution to journalReview articlepeer-review

Abstract

AbstractThe impressive success of large language models (LLMs) across a broad spectrum of NLP tasks has attracted considerable attention in both academia and industry. However, their inference incurs substantial computational and memory overhead, making it challenging to efficiently deploy them in latency-sensitive or resource-constrained environments. This survey provides a comprehensive overview of model-level optimization techniques developed to address these issues. Specifically, we examine efficient module designs for core Transformer components–including multi-head attention, feed-forward networks, normalization, and positional encoding–and summarize representative compression methods such as quantization, pruning, knowledge distillation, approximation, and parameter sharing. We categorize recent advances in each area, analyzing their algorithmic principles, empirical trade-offs, and impact on inference efficiency. Furthermore, we propose practical design strategies that effectively combine efficient modules with compression techniques to significantly reduce inference costs while meeting application-specific accuracy and latency requirements.

Original languageEnglish
Article number108900
JournalNeural Networks
Volume201
DOIs
StatePublished - Sep 2026

Keywords

  • Efficient inference
  • Large language models
  • Model-level optimization
  • Survey

Fingerprint

Dive into the research topics of 'Towards efficient language giants: A comprehensive survey on structural optimizations and compression techniques for large language models'. Together they form a unique fingerprint.

Cite this