PhonoFence: A Cross-Task Defense Framework for DeepFake via Phoneme-Level Adversarial Perturbations
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755061 ↗
摘要
Recent advances in deep neural networks and generative AI have enabled the creation of highly realistic synthetic speech, raising significant concerns about the misuse of deepfake in deception and fraud. This paper addresses the generalization challenge of active defense against cross-task audio deepfake under black-box conditions. We systematically analyze the common characteristics of state-of-the-art acoustic synthesis models and propose PhonoFence, an active defense framework that introduces fine-grained phoneme-level adversarial perturbations to prevent unauthorized synthesis. PhonoFence employs a dual-domain strategy to perturb both the time domain and frequency domain via an iterative cross-training framework, leveraging complementary acoustic features to enhance the generalization of perturbations. To further improve the transferability of perturbations, we ensemble speaker encoders with a novel Multi-Middle-Layer loss. Additionally, a psychoacoustic masking algorithm is employed to enhance the perceptual quality of protected speech and conceal perturbations. Extensive experiments on leading acoustic synthesis models demonstrate that PhonoFence reduces identity similarity and word error rate to 17.91% and 49.77%, respectively, achieving relative improvements of 7.97% and 9.67% over the best existing methods. To assess the effectiveness of our method in real-world, we test PhonoFence in commercial speaker recognition system, where it reduces deepfake attack success rates by 73.49%. Moreover, PhonoFence shows strong robustness against adaptive attacks involving compression, denoising, and re-recording.