Skip to main navigation Skip to search Skip to main content

Accelerated Training-Free Character-Consistent Text-to-Image Generation Framework

  • Doyun Kwon
  • , Geonoh Nam
  • , Mingyoo Song
  • , Hanul Kim
  • Seoul National University of Science and Technology (SNUST)

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We tackle character-consistent text-to-image (T2I) generation in a training-free setting. Shared attention with a subject mask enforces identity consistency across prompts but introduces substantial memory and latency overhead at inference. To mitigate this cost, we augment the stable diffusion baseline with two accelerations: token pruning, which removes redundant tokens, and adaptive guidance, which skips unnecessary computation during the diffusion process. Experiments on the ConsiStory+ benchmark show that our method outperforms recent state-of-the-art approaches in character-consistent T2I generation. Notably, it attains lower inference latency than both prior state-of-the-art methods and the SDXL baseline.

Original languageEnglish
Title of host publication2025 16th International Conference on Information and Communication Technology Convergence, ICTC 2025
PublisherIEEE Computer Society
Pages43-47
Number of pages5
ISBN (Electronic)9798331556785
DOIs
StatePublished - 2025
Event16th International Conference on Information and Communication Technology Convergence, ICTC 2025 - , Korea, Republic of
Duration: 14 Oct 202517 Oct 2025

Publication series

NameInternational Conference on ICT Convergence
ISSN (Print)2162-1233
ISSN (Electronic)2162-1241

Conference

Conference16th International Conference on Information and Communication Technology Convergence, ICTC 2025
Country/TerritoryKorea, Republic of
Period14/10/2517/10/25

Keywords

  • character-consistent image generation
  • efficient image generation
  • Text-to-image generation

Fingerprint

Dive into the research topics of 'Accelerated Training-Free Character-Consistent Text-to-Image Generation Framework'. Together they form a unique fingerprint.

Cite this