← 返回论文检索
ACM Multimedia 2025Experience: Multimedia Applications

SepVAMark: Deep Separable Visual-Audio Fusion Watermarking for Source Tracing and Deepfake Detection

Chuan Zhang 0003, Zihan Li, Zihao Xu, Xuhao Ren, Liehuang Zhu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755783 ↗

摘要

Visual-audio Deepfake has become increasingly prevalent in today's online environment. Passive detection methods, lacking preventive measures, struggle with detecting unknown forgery techniques, limiting their effectiveness. While proactive detection methods offer greater robustness, unimodal watermarking approaches remain vulnerable in visual-audio Deepfake scenarios, posing challenges to reliable forensics. To address these challenges, we propose a novel Separable Visual-Audio waterMark framework, called SepVAMark, for proactive Deepfake detection. SepVAMark incorporates a multi-layer perceptron-based mixer layer to fuse intra-modality and inter-modality features from both audio and visual data. We introduce the concept of separable visual-audio watermark, along with a bimodal robust extractor for traceability and two unimodal semi-robust extractors for Deepfake detection. This design ensures reliable copyright protection for source audio-video content while enabling authenticity verification for redistributed content. Experimental results on the FakeAVCeleb dataset demonstrate that SepVAMark effectively detects a wide range of advanced Deepfake manipulations, outperforming existing single-modal and multi-modal watermarking methods with superior robustness.