Skip to main navigation Skip to search Skip to main content

AE-Guide: Attention and Entropy Guided Visual Token Dropping for Accelerating High-Resolution Vision-Language Models

  • Seoul National University of Science and Technology (SNUST)

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

High-resolution vision-language models (VLMs) have emerged as powerful tools for multimodal tasks, but their reliance on numerous visual tokens from high-resolution inputs creates significant computational bottlenecks, especially in resource-constrained environments. Addressing this challenge, we introduce a novel, plug-and-play token-dropping method, engineered to operate within a predefined token budget. Our primary contribution is a refined visual token selection strategy that overcomes the limitations of traditional attention-guided methods. In contrast to traditional attention-guided approaches that leverage vision transformer (ViT) classification token attention for visual content measurement and budget allocation across image partitions, our approach incorporates the concept of token entropy to precisely identify the intrinsic information density and discriminative power of individual visual tokens. By capturing the inherent uncertainty within attention distributions, our method ensures that only the most salient and non-redundant visual tokens are preserved. We demonstrate that our proposed method effectively balances efficiency and accuracy across diverse VLM tasks. Empirical evaluations, particularly on the powerful LLaVA 1.6 Vicuna 13B model with a 20% token budget, show notable improvements. Compared to the attention-guided method, our approach achieved a 0.20% improvement in ScienceQA accuracy, a 9.57 point increase for MME, and a 0.09% gain in MMBench.

Original languageEnglish
Title of host publicationInternational SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331586423
DOIs
StatePublished - 2025
Event22nd International SoC Design Conference, ISOCC 2025 - Busan, Korea, Republic of
Duration: 15 Oct 202518 Oct 2025

Publication series

NameInternational SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers

Conference

Conference22nd International SoC Design Conference, ISOCC 2025
Country/TerritoryKorea, Republic of
CityBusan
Period15/10/2518/10/25

Keywords

  • Entropy-guided Compression
  • Vision-Language Model
  • Visual Token Dropping

Fingerprint

Dive into the research topics of 'AE-Guide: Attention and Entropy Guided Visual Token Dropping for Accelerating High-Resolution Vision-Language Models'. Together they form a unique fingerprint.

Cite this