UniEmotion: A Unified Framework for Multimodal Emotion Recognition with Iterative Consensus-based Training
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3762010 ↗
摘要
Traditional emotion recognition methods struggle with complex emotional dynamics including multi-emotion states, transitions, and contextual reasoning. While multimodal large language models demonstrate great potential for understanding such complex scene dynamics, they still face challenges in adapting to emotion recognition tasks. We propose UniEmotion, a unified framework that simultaneously addresses conventional categorical emotion recognition, open-vocabulary fine-grained emotion recognition, and descriptive emotion understanding. Our approach leverages an iterative consensus-based training pipeline where pseudo-labels and model parameters co-evolve, maximizing large models' utility while mitigating downstream limitations. The framework integrates a selector module that identifies high-quality samples through prediction variance analysis, coupled with a pseudo-labeling module employing consistency regularization and class-wise adaptive mapping. This dual mechanism reduces error accumulation during self-training while aligning open-vocabulary output with task-specific labels. Experimental results demonstrate the effectiveness of our framework, achieving state-of-the-art performance across all three tracks, including 1st place on the MER-SEMI track with a significant improvement of 11.97% over the best baseline, and 2nd place on the MER-DES track.