Towards Realistic Conversational Head Generation: A Comprehensive Framework for Lifelike Video Synthesis
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3581783.3612841 ↗
摘要
The Vivid Talking Head Video Generation track of the "ACM Multimedia ViCo 2023 Conversational Head Generation Challenge'' aims to generate realistic face-to-face conversation videos based on audio and reference images. However, the direct synthesis of reference videos from audio and reference images poses a significant challenge. In response, we propose a comprehensive method that combines audio-driven 3DMM parameter prediction with a rendering network to generate high-fidelity lip-sync talking head videos. In the first stage, we leverage the audio input to predict the 3DMM parameters, capturing essential facial expressions and head poses needed for realistic video generation. In the second stage, we employ a sophisticated rendering network designed to ensure accurate lip synchronization and natural facial expressions. This network enhances the visual quality and realism of the generated videos. We are delighted to announce that our method achieved first place in the talking head generation track of the challenge and was honored with the People's Selection Award. The source code for our method is accessible1, allowing others to build upon our work and drive further advancements in the field of conversational head generation.