← 返回论文检索
ACM Multimedia 2025Grand Challenges

Multiple Appropriate Facial Reaction Generation Based on Multi-View Transformation of Speaker Video

Jiajian Huang, Zitong Yu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3762246 ↗

摘要

Generating multiple appropriate facial reactions (MAFR) is essential for effective human-agent interaction. However, existing methods typically do not jointly model both local and global emotional cues, and neglect the temporal dynamics of facial expressions, leading to emotionally inconsistent and less natural reactions. In this work, we combine local and global emotional features to form a more comprehensive emotional representation. Our method further introduces motion-aware visual features that capture the dynamic evolution of facial expressions beyond static frames. By integrating both appearance and motion information within a structured generative framework, our approach enables more context-aware and temporally natural listener reactions. Experimental results demonstrate that our method outperforms existing approaches in both reaction diversity and appropriateness, which ranked first in the React 2025 challenge offline track.The implementation code can be accessed at:https://github.com/mtv-2025-react/mtv-2025-2025.git.