← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

DCount: Decoupled Spatial Perception and Attribute Discrimination for Referring Expression Counting

Ming Li 0083, Yupeng Hu 0003, Yinwei Wei, Hao Liu 0072, Haocong Wang, Weili Guan

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755872 ↗

摘要

Referring Expression Counting (REC) is an emerging task that aims to count specific objects in images based on textual phrases describing their attributes and categories. While current REC baselines inherit architectures from pre-trained open-vocabulary object detectors and demonstrate promising counting and localization capabilities, they overlook critical limitations in the original single-decoder design with shared object queries. This architectural constraint entangles the semantic and localization perception processes, hindering fine-grained understanding of attribute-aware visual features. To address these challenges, we propose DCount, a decoupled counting framework comprising two innovative components: a Decoupled Dual-Decoder (DDD) module and an Attribute Semantic Discriminator (ASD) module. The DDD module separates spatial perception tasks by employing distinct semantic and localization decoders with task-specific object queries, thereby enhancing the capture of discriminative visual features. Building upon the positional and semantic feedback from DDD, the ASD module introduces a two-stage filtering strategy to explicitly mine challenging hard negative attribute samples in the visual domain, while synergistically refining attribute discrimination across both modalities through contrastive learning in the textual domain. Our method achieves state-of-the-art results on both the REC and Zero-Shot Object Counting (ZSOC) benchmarks.