← 返回论文检索
EMNLP 2025emnlpfindings

Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Training

Fenghua Weng, Jian Lou, Jun Feng, Minlie Huang, Wenjie Wang

Sun Yat-Sen University · Huazhong University of Science and Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.findings-emnlp.735 ↗

摘要

Safety alignment is critical in pre-trained large language models (LLMs) to generate responses aligned with human values and refuse harmful queries. Unlike LLM, the current safety alignment of VLMs is often achieved with post-hoc safety fine-tuning. However, these methods are less effective to white-box attacks. To address this, we propose \textit{Adversary-aware DPO (ADPO)}, a novel training framework that explicitly considers adversary. \textit{Adversary-aware DPO (ADPO)} integrates adversarial training into DPO to enhance the safety alignment of VLMs under worst-case adversarial perturbations. \textit{ADPO} introduces two key components: (1) an adversarial-trained reference model that generates human-preferred responses under worst-case perturbations, and (2) an adversary-aware DPO loss that generates winner-loser pairs accounting for adversarial distortions. By combining these innovations, \textit{ADPO} ensures that VLMs remain robust and reliable even in the presence of sophisticated jailbreak attacks. Extensive experiments demonstrate that \textit{ADPO} outperforms baselines in terms of both safety alignment and general utility of VLMs.