Unsupervised Similarity-Fusion Transformer Hashing for Multimodal Retrieval
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754753 ↗
摘要
Unsupervised hashing is applied in large-scale multimodal retrieval by mapping original data from heterogeneous modalities into compact binary codes. Transformer-based retrieval augmented generation possesses significant advantages in retrieval accuracy and context-awareness, yet faces scalability challenges due to the computational overhead of dense embedding. Thus, the integration of hash learning and Transformer provides a feasible improvement scheme, which can achieve efficient retrieval preserving semantic association. This paper proposes a novel Unsupervised Similarity-Fusion Transformer Hashing for multimodal retrieval, denoted as USFTH. Initially, the modal fusion similarity matrix based on Gaussian kernel, sigmoid function, and Laplacian transformation is introduced to construct a discriminative similarity matrix, ensuring that semantic correlation among samples can be captured precisely. Then, cross-modal multiplex joint construction via Transformer-based attention mechanisms is designed, realizing effective integration of heterogeneous modalities in the similarity matrix through multi-path fusion. Furthermore, the consensus fusion strategy is proposed to ensure that hash codes generated under unsupervised conditions possess a uniform distribution and achieve accurate retrieval. In addition, comprehensive experiments on MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets demonstrate the superior performance of USFTH to state-of-the-art hashing approaches.