← 返回论文检索
ACM Multimedia 2025Content: Multimodal Fusion

CrosST: Cross Swin 4D Transformer for Multi-Modal Alzheimer's Detection

Hao Wang, Hanxiao Li, Li Xu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754980 ↗

摘要

The task of spatiotemporal dynamic modeling of multi-modal high-dimensional neuroimaging data presents a significant challenge in the field of neuroscience. Recent works often integrate attention mechanisms for hierarchical modeling, but the gradual extraction of spatiotemporal features leads to feature isolation. Moreover, attention-based fusion mechanisms (such as cross attention) tend to focus on learning the self-similarity between different modalities, lacking sufficient exploration of the complementary information across modalities. To address these challenges, we propose the cross swin 4D transformer (CrosST), which can efficiently learn the spatiotemporal patterns of multi-modal high-dimensional neuroimaging data in an end-to-end manner. The unique diffusion cross attention fusion mechanism of CrosST connects features from different modalities through a diffusion strategy during the attention computation, enabling the transfer of differential information between modalities and achieving deep fusion of multi-modal coupled features. Additionally, a voxel interaction strategy is employed to alleviate the computational burden during the fusion process. Furthermore, CrosST utilizes a 4D shifted window technique to effectively combine local and global information, and introduces the innovative 4D-Mamba algorithm to enhance computational efficiency. We validate the model using a large-scale Alzheimer's disease dataset and design a multi-granularity cognitive stage task for evaluation. The results demonstrate the effectiveness of CrosST.