← 返回论文检索
ACM Multimedia 2025Engagement: Emotional and Social Signals

MoCERNet: A Modality-Complete Modeling Framework for Emotion Recognition in Physiological Signals under Imperfect Modal Matching

Tianzuo Xin, Jing Wang 0060, Xiyuan Jin, Xiaojun Ning 0001, Zhiyang Feng, Youfang Lin

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755354 ↗

摘要

Emotion recognition based on multimodal physiological signals is playing an increasingly important role in areas such as human-computer interaction and disease diagnosis, attracting growing attention from the research community. Current studies primarily focus on emotion recognition under unified data collection paradigms, overlooking the prevalent issue of imperfect modality matching in real-world scenarios. In particular, existing methods fail to effectively utilize these mismatched modalities, leading to incomplete emotional representations. This limits the model's ability to accurately capture the multidimensional semantic features of emotions, thereby constraining its effectiveness and applicability in practical settings. To address this challenge, we propose MoCERNet. At the modality level, it first reduces the domain gap among matched modalities and then aligns mismatched modalities in a semantics-aware manner, guided by the matched ones. At the decision level, it further mitigates the global distribution discrepancies to achieve a more complete emotional representation. In addition, we design a Nervous System Functional Structure Transformer (NFSformer) that enables the model to focus on the correlation between different brain regions and peripheral physiological signals under various emotional states, thereby enhancing its capacity to model complex emotional processes. Experiments on three multimodal emotion datasets demonstrate that MoCERNet outperforms state-of-the-art baselines under imperfect modality matching scenarios.