Contextually-Guided State Space Fusion for Misaligned Multi-Spectral Object Detection
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754550 ↗
摘要
Multi-modal feature fusion under conditions of image misalignment remains a significant challenge in multispectral object detection. Existing approaches predominantly rely on cross-attention mechanisms; however, when local features are sparse, inadequate feature capture hinders accurate alignment and results in distorted fusion outcomes. To address this problem, we propose a novel multispectral fusion detection network, CSSFDet, which leverages the intrinsic correlations among image regions in visual recognition to dynamically enhance local features via global semantic constraints during the fusion process. Specifically, we introduce a Contextual Region Feature Fusion Module (CRFM) that regulates the fusion process through a selective state-space formulation, adaptively incorporating surrounding context to compensate for local feature degradation caused by misalignment. Moreover, we design a Complementary Enhancement Module (CoE) to ensure both distinctiveness and completeness of modality-specific features. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across multiple datasets, attaining 84.1% mAP50 on the DroneVehicle dataset-a 20% improvement over the baseline. It also shows strong performance on the misaligned CVC-14 dataset, and sensitivity analysis on data shifts further underscores its robustness to misalignment.