PureProof: Diffusion-Resistant Black-box Targeted Attack on Large Vision-Language Models
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Large Vision-Language Models (VLMs) are increasingly deployed across diverse applications, such as AI agents, yet remain vulnerable to targeted adversarial attacks. However, the practical robustness of such attacks often remains unclear with limited evaluation under defenses. Diffusion-based purification (DBP), a widely adopted black-box defense, effectively blocks current attacks by removing adversarial perturbations via generative diffusion. Prior DBP evasion attacks target white-box image classifiers and are ill-suited to VLMs, incurring high computational costs and gradient instability when adapted. In this paper, we present PureProof, a black-box targeted attack on VLMs resilient to DBP. It consists of three core components. Stochastic Reverse Alignment (SRA) guides adversarial optimization via single-step reverse prediction, avoiding costly full-trajectory backpropagation. Adaptive Re-noising Augmentation (ARA) mitigates diffusion stochasticity through timestep-adaptive re-noising. Self-Consistency Regularization (SCR) stabilizes optimization by promoting local temporal coherence. Extensive experiments on open-source and commercial VLMs show that PureProof consistently outperforms prior attacks against DBP and achieves strong noise resilience, revealing critical vulnerabilities in VLMs and offering broader safety implications for real-world VLM deployments, including emerging agentic settings.