Entity-Level Alignment with Prompt-Guided Adapter for Remote Sensing Image-Text Retrieval
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755374 ↗
摘要
Remote Sensing Image-Text Retrieval (RSITR) is a fundamental task in the remote sensing (RS) field and has seen significant progress in recent years. However, existing methods often overlook explicit attention to semantic entities in RS scenes, limiting their capabilities in fine-grained semantic modeling and cross-modal matching, thereby hindering retrieval performance. To address these limitations, we propose a novel framework, Entity-level Alignment with Prompt-guided Adapter (EAPA), which enhances retrieval performance by explicitly perceiving, embedding, and aligning semantic entities in RS images and texts. Built upon the Contrastive Language-Image Pretraining (CLIP) model, EAPA comprises three key modules: the Prompt-guided Attention Adapter (PAA) module, the Pseudo-label-supervised Entity Embedding (PEE) module, and the Cross-modal Entity-level Semantic Alignment (CESA) module. Specifically, PAA freezes the CLIP backbone and introduces learnable prompt vectors to capture RS-specific entity-level semantic knowledge, guiding attention distribution and enhancing semantic representations. To obtain cross-modal consistent entity-level representations, PEE employs an entity query-based encoder to extract entity embeddings of both images and texts, and uses pseudo semantic labels as supervision to ensure that each embedding corresponds to a unique and well-defined semantic category. Based on this, CESA performs one-to-one alignment of cross-modal entity embeddings that correspond to the same semantic category, effectively avoiding mismatches and enhancing fine-grained alignment. Extensive experiments on the RSICD and RSITMD datasets demonstrate that EAPA outperforms state-of-the-art methods across multiple metrics, validating the effectiveness of each module in enhancing fine-grained semantic modeling and cross-modal matching.