TY - GEN
T1 - SGGA
T2 - 25th International Conference on Control, Automation and Systems, ICCAS 2025
AU - Jeong, Dayena
AU - Heo, Dongwook
AU - Ahn, Seonghyeok
AU - Choi, Jonggeun
AU - Choi, Sunglok
N1 - Publisher Copyright:
© 2025 ICROS.
PY - 2025
Y1 - 2025
N2 - Object detection in disaster images is often difficult because of class imbalance and a lack of labeled samples. To solve this problem, we introduce Semantic-Guided Generative Augmentation (SGGA), a new method that uses semantic masks to generate more samples for the rare classes. SGGA creates new images by changing clean road areas into Road-Blocked areas using mask-based sampling and prompt-guided inpainting, making sure the new objects appear in the right places. We filter the new images using CLIP similarity and LPIPS distance, ensuring high semantic and visual quality. Experiments on the RescueNet dataset show that SGGA improves Road-Blocked detection by +26.2 % mAP @ 0.5 and +29.7 % recall, beating other augmentation methods. Furthermore, t-SNE analysis confirms strong semantic alignment between real and SGGA-generated images. SGGA offers significant advantages, including spatial precision, contextual realism, and low annotation overhead, making it particularly suitable for practical deployment in disaster scenarios and other domains where spatial priors are available.
AB - Object detection in disaster images is often difficult because of class imbalance and a lack of labeled samples. To solve this problem, we introduce Semantic-Guided Generative Augmentation (SGGA), a new method that uses semantic masks to generate more samples for the rare classes. SGGA creates new images by changing clean road areas into Road-Blocked areas using mask-based sampling and prompt-guided inpainting, making sure the new objects appear in the right places. We filter the new images using CLIP similarity and LPIPS distance, ensuring high semantic and visual quality. Experiments on the RescueNet dataset show that SGGA improves Road-Blocked detection by +26.2 % mAP @ 0.5 and +29.7 % recall, beating other augmentation methods. Furthermore, t-SNE analysis confirms strong semantic alignment between real and SGGA-generated images. SGGA offers significant advantages, including spatial precision, contextual realism, and low annotation overhead, making it particularly suitable for practical deployment in disaster scenarios and other domains where spatial priors are available.
KW - Class imbalance
KW - disaster imagery
KW - generative data augmentation
KW - image data augmentation
KW - rare object detection
KW - semantic-guided inpainting
KW - stable diffusion
UR - https://www.scopus.com/pages/publications/105031893041
U2 - 10.23919/ICCAS66577.2025.11301360
DO - 10.23919/ICCAS66577.2025.11301360
M3 - Conference contribution
AN - SCOPUS:105031893041
T3 - International Conference on Control, Automation and Systems
SP - 1661
EP - 1664
BT - 2025 25th International Conference on Control, Automation and Systems, ICCAS 2025
PB - IEEE Computer Society
Y2 - 4 November 2025 through 7 November 2025
ER -