TY - GEN
T1 - AE-Guide
T2 - 22nd International SoC Design Conference, ISOCC 2025
AU - Lee, Gilha
AU - Kim, Hyun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - High-resolution vision-language models (VLMs) have emerged as powerful tools for multimodal tasks, but their reliance on numerous visual tokens from high-resolution inputs creates significant computational bottlenecks, especially in resource-constrained environments. Addressing this challenge, we introduce a novel, plug-and-play token-dropping method, engineered to operate within a predefined token budget. Our primary contribution is a refined visual token selection strategy that overcomes the limitations of traditional attention-guided methods. In contrast to traditional attention-guided approaches that leverage vision transformer (ViT) classification token attention for visual content measurement and budget allocation across image partitions, our approach incorporates the concept of token entropy to precisely identify the intrinsic information density and discriminative power of individual visual tokens. By capturing the inherent uncertainty within attention distributions, our method ensures that only the most salient and non-redundant visual tokens are preserved. We demonstrate that our proposed method effectively balances efficiency and accuracy across diverse VLM tasks. Empirical evaluations, particularly on the powerful LLaVA 1.6 Vicuna 13B model with a 20% token budget, show notable improvements. Compared to the attention-guided method, our approach achieved a 0.20% improvement in ScienceQA accuracy, a 9.57 point increase for MME, and a 0.09% gain in MMBench.
AB - High-resolution vision-language models (VLMs) have emerged as powerful tools for multimodal tasks, but their reliance on numerous visual tokens from high-resolution inputs creates significant computational bottlenecks, especially in resource-constrained environments. Addressing this challenge, we introduce a novel, plug-and-play token-dropping method, engineered to operate within a predefined token budget. Our primary contribution is a refined visual token selection strategy that overcomes the limitations of traditional attention-guided methods. In contrast to traditional attention-guided approaches that leverage vision transformer (ViT) classification token attention for visual content measurement and budget allocation across image partitions, our approach incorporates the concept of token entropy to precisely identify the intrinsic information density and discriminative power of individual visual tokens. By capturing the inherent uncertainty within attention distributions, our method ensures that only the most salient and non-redundant visual tokens are preserved. We demonstrate that our proposed method effectively balances efficiency and accuracy across diverse VLM tasks. Empirical evaluations, particularly on the powerful LLaVA 1.6 Vicuna 13B model with a 20% token budget, show notable improvements. Compared to the attention-guided method, our approach achieved a 0.20% improvement in ScienceQA accuracy, a 9.57 point increase for MME, and a 0.09% gain in MMBench.
KW - Entropy-guided Compression
KW - Vision-Language Model
KW - Visual Token Dropping
UR - https://www.scopus.com/pages/publications/105033154577
U2 - 10.1109/ISOCC66390.2025.11329950
DO - 10.1109/ISOCC66390.2025.11329950
M3 - Conference contribution
AN - SCOPUS:105033154577
T3 - International SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers
BT - International SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 15 October 2025 through 18 October 2025
ER -