← 返回论文检索
ACM Multimedia 2025Generative AI: Generative Multimedia

Image Retargeting based on Text Region Awareness

Gang Pan 0002, Meihua Liu, Lei Zhou 0031, Jiahao Wang, Di Sun 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755414 ↗

摘要

Image retargeting (IR) with text regions is a challenging yet underexplored task that focuses on resizing an image's aspect ratio while preserving both semantic objects and the legibility of textual content. This task introduces three primary challenges, which can be summarized as follows: (1) the distinct probability distributions between text and non-text regions in images; (2) the lack of dedicated mechanisms in existing IR methods for handling text regions, often leading to text distortion or blurring; (3) the absence of paired datasets specifically designed for IR tasks involving text regions. To tackle these challenges, we propose SSIR, a unified framework that reformulates IR as a joint Semantic Segmentation and Image Retargeting (SS-IR) task, leveraging an attention mechanism to bridge these components. Specifically, we first employ a semantic segmentation sub-network that extracts text region features using text segmentation techniques to improve text-awareness in retargeting tasks. Then, we integrate text features into the image's visual representation through an attention-driven module designed to preserve both textual and semantic content during retargeting. Finally, we address the absence of paired datasets with an unsupervised learning paradigm based on a Cycle-IR framework, which employs cyclic consistency reconstruction, enabling effective learning without the need for paired training data. Experimental results show that the SSIR algorithm effectively preserves text information and delivers high-quality visual retargeting results.