← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Domain-aware Visual Context Prompt for Multi-Source Domain Adaptation

Yuwu Lu, Haoyu Huang, Xue Hu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754547 ↗

摘要

Leveraging pre-trained Vision-Language Models (VLMs) for downstream tasks has gained significant attention recently, particularly in Multi-Source Domain Adaptation (MSDA). However, most existing VLMs-based MSDA approaches rely on domain-specific text prompts, which struggle to capture domain-invariant representations. In addition, multiple domains in MSDA introduce significant distribution discrepancies, complicating the design of effective text prompts. To address these challenges, we propose a Domain-aware Visual Context Prompt (DVCP) method, which leverages domain-level features to bridge the domain gaps. Specifically, we design domain-aware text prompts (DTP) module that maps global visual information into the textual prompt embedding space, creating trainable text prompts that incorporate domain-level visual information. Then, we construct a domain-aware visual tuning (DVT) module that collaboratively leverages domain-level and instance-level features to align distributions across multiple domains. Extensive experiments conducted on four popular MSDA benchmarks including Office31, ImageCLEF-DA, Office-Home, and DomainNet, demonstrate the superiority of the proposed method.