CIA: Class- and Instance-aware Adaptation for Vision-Language Models
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754846 ↗
摘要
Few-shot parameter-efficient tuning methods demonstrate promising potential for Vision-Language (V-L) models in downstream tasks. However, existing approaches primarily focus on class-level alignment between image and text features, overlooking crucial instance-specific semantic information. This limitation leads to suboptimal performance on challenging tasks and restricted generalization capability to unseen data. To address these issues, we propose Class- and Instance-aware Adaptation (CIA), a novel framework that simultaneously optimizes both class-level and instance-level alignments. Specifically, CIA introduces a novel instance encoder that leverages cross-modal self-attention to generate instance-specific text features, accompanied by a carefully designed regularization mechanism to maintain consistency between class-level and instance-level representations. Extensive experiments across 15 benchmark datasets demonstrate that CIA significantly improves the downstream adaptation of V-L models.