← 返回论文检索
ICML 2026PosterAccept (regular)

Locate then Correct: Debiasing Attention Heads in CLIP

Wei Yeo, Rui Mao, Moloud Abdar, Ranjan Satapathy, Erik Cambria

School of Computer Science and Engineering, Nanyang Technological University · Nanyang Technological University · Deakin University · A*STAR

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Multimodal models like CLIP have gained significant attention due to their remarkable zero-shot performance across various tasks. However, studies have revealed that CLIP can inadvertently learn spurious associations between target variables and confounding factors. To address this, we introduce \textsc{Locate-Then-Correct} (LTC), a contrastive framework that identifies spurious attention heads in Vision Transformers via mechanistic insights and mitigates them through targeted ablation. Furthermore, LTC identifies salient, task-relevant attention heads, enabling the integration of discriminative features through orthogonal projection to improve classification performance. We evaluate LTC on benchmarks with inherent background and gender biases, achieving over a > 50% gain in worst-group accuracy compared to non-training post-hoc baselines. Additionally, we visualize the representation of selected heads and find that the presented interpretation corroborates our contrastive mechanism for identifying both spurious and salient attention heads.