← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person Retrieval

Yifei Deng, Chenglong Li 0002, Futian Wang, Jin Tang 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754721 ↗

摘要

Existing Text-Image Person Retrieval (TIPR) methods have made substantial progress in modeling cross-modal associations via contrastive learning frameworks, but usually ignore the fine-grained differences in semantic relevance among different samples, which limits retrieval accuracy. To address this problem, we propose a novel Hierarchical Cross-modal Association framework HCA, which leverages the intra-modal fine-grained semantic relations distilled by single-modal pretrained models to constrain hierarchical cross-modal association between image and text modalities, for accurate TIPR. Specifically, to model hierarchical cross-modal semantic relationships, we propose a Hierarchical Relevance Matching (HRM) module. It partitions the matching strength of image-text pairs by jointly considering identity labels and cross-modal similarity, collaborating with unimodal similarity to construct a hierarchical relevance distribution that serves as a soft supervision signal. HRM not only helps the model better capture varying levels of semantic consistency between image-text pairs but also enhances the overall accuracy of cross-modal association learning. To enhance the ability to capture fine-grained cross-modal semantic relationships, we introduce an Image-guided Ambiguous text Token Modeling (IATM) module. It replaces original tokens with semantically ambiguous ones and leverages image guidance to detect and correct these tokens. This process further improves the fine-grained semantic alignment between images and texts. Experimental results demonstrate that HCA achieves new state-of-the-art performance across multiple datasets, thoroughly validating its effectiveness and advancement in cross-modal retrieval tasks.