← 返回论文检索
ACM Multimedia 2025Generative AI: Generative Multimedia

REA-Listener: Real-Time Listening Head Generation with Dynamic Emotion Modeling and Flexible Modality Adaptation

Sizhe Zhao, Chenyang Wang 0002, Weiyu Zhao, Zonglin Li 0004, Ming Li 0042, Shengping Zhang

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755093 ↗

摘要

Listening head generation aims to synthesize realistic and responsive non-verbal listener head motions that respond to speakers in conversational scenarios. Existing methods typically rely on fixed audio-visual input modalities and predefined emotion labels, limiting their adaptability and expressiveness in real-world scenarios. In this paper, we propose a novel real-time framework, REA-Listener, to generate high-fidelity listening head videos with flexible modality adaptation and dynamic emotion modeling. Specifically, we first propose a Modality-Adaptive Mixture of Experts (MA-MoE) module to encode arbitrary combinations of speaker audio and visual signals into a unified embedding space, ensuring robustness under partial modality conditions. To further enhance the temporal consistency of listener emotion, we present a lightweight emotional head dynamics generator with a multi-modal emotion predictor, which infers listener emotions dynamically from speaker context alongside head motion coefficient prediction. Finally, we employ a 3D-aware renderer based on 3D Gaussian Splatting to produce high-quality listener head videos in real time. With these components, our approach achieves efficient head motion generation at 30fps on a single NVIDIA RTX 3090 GPU, supporting real-time interaction. Extensive evaluations and applications demonstrate that our method outperforms state-of-the-art methods in listening head generation.