← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

AFFIR: Dual-Modal Attention Feature Fusion for Scene Text Image Retargeting

Gang Pan 0002, Liming Pan, Hongze Mi, Rongyu Xiong, Jiahao Wang, Di Sun 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755838 ↗

摘要

Image retargeting technique aims to adjust and reorganize the content of original images to fit different display sizes and visual requirements. Text elements frequently appear in real-world images and play a crucial role in conveying information. Existing algorithms often treat the image as a whole during retargeting, neglecting the unique features of textual content. This oversight results in missing textual information or distorted character structures, ultimately failing to effectively preserve the integrity of text regions, thereby affecting both the efficiency of information transmission and visual quality of the final image. To address the aforementioned issues, we start from the perception of textual content, which guides retargeted image generation through the fusion of attention features. Specifically, a Transformer-based model is employed for the image retargeting tasks in this study. Text and image features are extracted separately, accompanied by a dual-modal feature fusion strategy, which integrates text and image features through attention maps generated. The training process adopts a cyclic training strategy, where the retargeted results are fed back into the model in reverse. This approach is applicable to retargeting images of various sizes, ensuring that detailed information from both text and image content is accurately preserved. Extensive evaluations on benchmark datasets demonstrate that our method significantly outperforms existing techniques in maintaining both textual clarity and overall visual quality, making it a promising solution for advanced multimedia applications in computer science.