WhiADD: Semantic-Acoustic Fusion for Robust Audio Deepfake Detection
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755594 ↗
摘要
This paper addresses the critical challenge of detecting codec-based audio deepfakes in multilingual and dynamically evolving adversarial scenarios. While existing detection systems exhibit performance degradation against codec-generated forgeries and unseen linguistic environments, we propose a novel audio deepfake detection framework ''WhiADD'' enhanced by semantic-acoustic fusion and cross-modal generalization. Our methodology introduces three key innovations: (1) The Union CodecFake (UCF) dataset, synthesized by extending the CodecFake generation pipeline to the multilingual Common Voice corpus, significantly expands acoustic diversity with 1.9M samples across varied phonetic, channel, and codec manipulation patterns. (2) A semantic-prompted Whisper architecture that integrates full-transcript linguistic constraints into decoder fine-tuning, enabling detection of semantic inconsistencies. (3) A gated cross-attention mechanism that dynamically fuses multi-source audio features with the proposed model's frozen encoder outputs, enhancing artifact detection through adaptive attention to pre-trained representations. Extensive experiments demonstrate state-of-the-art performance, achieving 0.55% EER on UCF testing data and less than 3% EER in zero-shot cross-lingual detection (German, French, Italian). The framework reduces false negatives by up to 24% compared to conventional models through improved semantic-acoustic alignment. These advancements establish a robust paradigm for combating evolving codec-based forgeries, bridging the critical gap between acoustic feature engineering and semantic coherence analysis in audio forensics.