Skip to main navigation Skip to search Skip to main content

DQ-LUT: Dyadic Quantization-Aware LUT-based Approximation for Non-linear Functions in VisionTransformers

  • Seoul National University of Science and Technology (SNUST)

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Vision transformers (ViTs) deliver state-of-the-art accuracy on vision tasks, and their adoption is expanding to on-device AI in privacy- and latency-sensitive domains (e.g., defense, medical). Yet most prior optimizations target matrix multiplication; once those are accelerated, the runtime bottleneck shifts to nonlinear functions, undermining end-to-end hardware gains. Conventional approximation of these functions often degrades accuracy and requires fine-tuning - impractical where training data are restricted. We propose DQ-LUT, a dyadic quantization-aware, LUT-based approximation that preserves model performance without fine-tuning. Our method minimizes quantization error by factorizing each scale into a power-of-two (dyadic) component and a residual shift, then fusing both with piecewise-linear (PWL) parameters. This enables the entire approximation to execute on a pure INT8 MAC datapath, eliminating floating-point and costly divisions. To handle wide input ranges, we further introduce a hardware-friendly Multi-Range Input Scaling (MRIS) module tailored to nonlinear operators. On INT8-quantized ViT models, DQ-LUT limits accuracy loss to ≤ 1.418% without any fine-tuning. FPGA synthesis on a Xilinx ZU9EG shows the proposed MRIS reduces LUT usage by 34.9% compared to a conventional MRIS design, confirming superior hardware efficiency. These results demonstrate a practical path to fine-tuning-free, accelerator-friendly nonlinear approximation for on-device ViT deployment.

Original languageEnglish
Title of host publication2025 IEEE/IEIE International Conference on Consumer Electronics-Asia, ICCE-Asia 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331574024
DOIs
StatePublished - 2025
Event2025 IEEE/IEIE International Conference on Consumer Electronics-Asia, ICCE-Asia 2025 - Busan, Korea, Republic of
Duration: 27 Oct 202529 Oct 2025

Publication series

Name2025 IEEE/IEIE International Conference on Consumer Electronics-Asia, ICCE-Asia 2025

Conference

Conference2025 IEEE/IEIE International Conference on Consumer Electronics-Asia, ICCE-Asia 2025
Country/TerritoryKorea, Republic of
CityBusan
Period27/10/2529/10/25

Keywords

  • Edge computing
  • Fieldprogrammable gate arrays (FPGAs)
  • Nonlinear function approximation
  • Quantization
  • Vision Transformers (ViT)

Fingerprint

Dive into the research topics of 'DQ-LUT: Dyadic Quantization-Aware LUT-based Approximation for Non-linear Functions in VisionTransformers'. Together they form a unique fingerprint.

Cite this