DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Multimodal tiny object detection plays a critical role in real-world applications, yet remains highly challenging due to weak target representations and complex cross-modal interference. Existing frequency-domain methods for tiny object detection are still largely limited to the visible modality and overlook complementary cross-modal frequency cues in multimodal scenes. In this paper, we investigate cross-modal frequency learning for RGBT tiny object detection. Through frequency characteristic analysis, we find that tiny objects in both RGB and infrared modalities contain richer mid- and high-frequency components as object size decreases. Motivated by this observation, we propose a Dynamic Frequency-Decoupled Cross-Modal Learning Transformer (DyFCLT). Specifically, DyFCLT introduces a Dynamic Frequency-Band Decoupled Cross-Modal Attention (DFCA) mechanism to perform fine-grained cross-modal interaction across dynamic frequency sub-bands, and a Selective Smoothing Enhancement (SSE) module to suppress background noise and enhance foreground responses during multi-scale fusion. Extensive experiments on two RGBT tiny object detection benchmarks and one general-scale benchmark demonstrate that DyFCLT achieves state-of-the-art performance with strong generalization across different scales and scenes.