← 返回论文检索
ACM Multimedia 2024Poster Session 1

Overcoming the Pitfalls of Vision-Language Model for Image-Text Retrieval

Feifei Zhang 0001, Sijia Qu, Fan Shi 0001, Changsheng Xu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3664647.3680591 ↗

摘要

This work tackles the persistent challenge of image-text retrieval, a key problem at the intersection of computer vision and natural language processing. Despite significant advancements facilitated by large-scale Contrastive Language-Image Pretraining (CLIP) models, we found that existing methods fall short in bridging the fine-grained semantic gap between visual and textual representations. To address the above pitfalls, we propose a model called Local and Generative-driven Modality Gap Correction (LG-MGC), which devotes to simultaneously enhancing representation learning and alleviating the modality gap in cross-modal retrieval. The proposed model consists of two main components: a local-driven semantic completion module, which complements specific local context information that is overlooked by traditional models within global features, and a generative-driven semantic translation module, which leverages generated features as a bridge to mitigate modality gap. Our model not only tackles the granularity of semantic correspondence and improves the performance of existing methods without requiring additional trainable parameters, but is also designed to be plug-and-play, allowing for easy integration into existing retrieval models without altering their architectures. Extensive experiments demonstrate the effectiveness of LG-MGC by achieving consistent state-of-the-art performance over strong baselines.