Mitigating the Evolving Semantic Entanglement in Continual Learning of Vision-Language Models
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755498 ↗
摘要
Multi-domain Task Incremental Learning (MTIL) aims to continuously acquire knowledge from diverse domains while maintaining generalization capability. Recent works have demonstrated promising results by leveraging Vision-Language Models (VLMs) for continual learning. However, we identify a critical issue within this paradigm, termed the Evolving Semantic Entanglement. Specifically, VLMs tend to produce highly similar text features for semantically related categories, resulting in hard feature alignment and interference between related categories. This problem becomes increasingly pronounced as the category space expands. In this paper, we present a novel Dual-granularity Prompt Learning (DuPLe) framework to address this challenge. Our approach enhances text feature discriminability by leveraging complementary two-level prompts: category-level global prompts for holistic semantic concepts and attribute-level local prompts for fine-grained visual patterns. We further apply a Task Assignment-free Inference strategy that eliminates explicit task identification, simplifying the inference process and enabling extension to unseen categories. Extensive experiments on 11 diverse domains under MTIL and X-TAIL settings demonstrate that our method significantly mitigates the entanglement issue and outperforms previous state-of-the-art approaches.