Skip to main navigation Skip to search Skip to main content

GDE: Grid-based Diversity Enhancement for Efficient Vision Token Selection in VLMs

  • Seoul National University of Science and Technology (SNUST)

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent vision-language models (VLMs) have achieved impressive performance across various tasks but suffer from high computational cost due to the large number of visual tokens. While diversity-based token selection methods effectively reduce redundancy, they often lose spatial coverage under tight token budgets. To address this issue, we propose grid-based diversity enhancement (GDE), a lightweight module that enforces balanced spatial coverage by ensuring that each grid region contributes at least one representative token. $GDE$ can be seamlessly integrated into existing diversity-based frameworks with minimal overhead. Experiments on image understanding benchmarks using LLaVA-1.5-7B demonstrate that $GDE$ consistently improves performance, achieving $97.4 \%$ of the vanilla model's accuracy with only 128 tokens. These results show that $GDE$ effectively preserves global visual context and achieves a better trade-off between efficiency and accuracy under extreme compression.

Original languageEnglish
Title of host publication2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331580773
DOIs
StatePublished - 2026
Event2026 International Conference on Electronics, Information, and Communication, ICEIC 2026 - Macau, China
Duration: 18 Jan 202621 Jan 2026

Publication series

Name2026 International Conference on Electronics, Information, and Communication, ICEIC 2026

Conference

Conference2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
Country/TerritoryChina
CityMacau
Period18/01/2621/01/26

Keywords

  • multimodal learning
  • token dropping
  • token selection
  • Vision language model

Fingerprint

Dive into the research topics of 'GDE: Grid-based Diversity Enhancement for Efficient Vision Token Selection in VLMs'. Together they form a unique fingerprint.

Cite this