DiffuSeg: Diffusion-Enhanced Cross-Modal Semantic Segmentation for RGB-D
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755706 ↗
摘要
Diffusion models have recently shown strong capabilities in image generation. This paper investigates their potential for semantic segmentation, with a focus on RGB-D tasks that demand precise pixel-level predictions. In particular, we delve into the intermediate activations generated during the reverse Markov step of diffusion process, discovering that these activations can effectively capture the semantic information of an input image, making them outstanding representations for addressing segmentation challenges. This paper proposes Diffusion-Enhanced Multi-Modal Segmenter (DiffuSeg), which innovatively combines RGB features with those generated by an additional diffusion model, facilitating the extraction of comprehensive and nuanced semantic features. Furthermore, we propose the Cross Attention-and-Aggregation Module (CAAM), which not only fosters long-range interactions between RGB and diffusion-derived features but also recalibrates both feature sets before integration, enhancing multi-modal synergy. Additionally, our model incorporates a Dynamic Cascade Kernel (DCK) architecture that exploits local and intricate multi-scale geometric details. As a part of DCK, the Spatial Interaction Module (SIM) dynamically encodes spatial information by establishing pixel-level correlations, thereby enhancing the spatial feature representation capacity. Extensive experiments on two benchmark datasets demonstrate the strong capability of DiffuSeg in handling challenging semantic segmentation tasks.