← 返回论文检索
ACM Multimedia 2025Content: Multimodal Fusion

Dynamic Optimization Noisy Cross-Modal Hashing

Zebing Yao, Hao Fu 0020, Yuanhang Yang, Guanghua Gu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755684 ↗

摘要

Cross-Modal Hashing (CMH) has gained significant attention for its ability to learn semantic category discrimination and enable efficient retrieval. However, in practical applications, the massive amounts of multi-modal data collected from the internet often contain coarse annotations, which inevitably introduce noisy labels and degrade retrieval performance. To address this challenge, this paper proposes a dynamic optimization-based training framework, namely Dynamic Optimization Noisy Cross-Modal Hashing (DONCMH). Firstly, to alleviate the issue of overfitting to noisy labels during training, we propose a novel regularization-based noise-robust strategy that updates the target distribution with momentum to optimize clustering learning, thus avoiding over-emphasizing noisy samples. Secondly, to more accurately select high-quality training samples, we introduce ClusterOT, a novel Optimal Transport formulation explicitly tailored for Noisy Cross-Modal Hashing (NCMH), which integrates center representation learning and cross-modal alignment into a unified structure. By leveraging the spatial distribution of samples, ClusterOT effectively mitigates distribution imbalances inherent in center representation learning, thereby significantly improving the model's robustness to noisy label predictions. Finally, a robust feature learning module is employed to enhance the extraction of informative and discriminative representations from both modalities. Extensive experiments conducted on four widely used benchmark datasets demonstrate that the proposed method effectively mitigates the impact of noisy labels and significantly improves cross-modal retrieval performance.