Audio-Visual Asynchrony Mitigation: Cross-Modal Alignment and Feature Reconstruction for Deepfake Detection
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755455 ↗
摘要
The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has enabled deepfake videos to evolve from unimodal generation to audio-visual forgeries. Existing multimodal deepfake detection methods primarily rely on capturing correlations between audio-visual modalities to improve detection performance. However, in real-world scenarios, network jitter often leads to audio-visual asynchrony, disrupting inter-modal associations and limiting the effectiveness of these methods. To address this issue, we propose a deepfake detection method specifically designed for audio-visual asynchrony scenarios. First, based on the theory of open balls in metric space, we analyze the variation mechanism of joint features in both audio-visual synchrony and asynchrony scenarios, revealing the impact of audio-visual asynchrony on detection performance. Second, we design a multimodal subspace representation module to mitigate inconsistencies in feature distributions and representation heterogeneity between modalities. We then formulate audio-visual feature alignment as an integer linear programming task and employ the Hungarian algorithm to reconstruct missing inter-modal associations. Finally, we introduce a self-supervised masked reconstruction mechanism to reconstruct missing features and construct the joint correlation matrix to measure cross-modal dependencies, enhancing the robustness of detection. Extensive experiments demonstrate that our method outperforms baselines in audio-visual asynchrony scenarios and exhibits robustness against unknown disturbances.