← 返回论文检索
ACM Multimedia 2025Content: Multimodal Fusion

A Motion is Worth a Hybrid Sentence: Taming Language Model for Unified Motion Generation by Fine-grained Planning

Ronghui Li, Lingxiao Han, Shi Shu, Yueyao Liu, Yukang Lin, Yue Ma 0016, Jie Guo, Ziwei Liu 0002, Xiu Li 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755164 ↗

摘要

Existing LLM-based motion models fail to fully leverage large models' planning capabilities for motion-related tasks, exhibiting poor generalization, limited text-motion alignment, and an inability to perform multimodal condition joint driven motion generation. We argue that these issues arise from the modality gap and the highly coupled nature of motion tokens. To address this, we proposed the hybrid motion sentence, which is consistant of fine-grained motion decription and atomic body-part motion token that can bridge the gap between motion and text. To obtain a large corpus of hybrid motion sentences, we introduced a novel motion-to-text generation method that combines atomic motion operators with GPT-4o, resulting in 68.2 million fine-grained textual descriptions across diverse modalities. To reconstruct high-quality motion from hybrid sentences and make better motion-text alignment, we introduce Semantic-Aware Decoupled Motion Tokenization. Furthermore, we propose MotionUPG based on LLaMA, leveraging MotionWords dataset for both pretraining and instruction tuning. Our method achieves strong fine-grained text-motion alignment, impressive zero-shot motion generation, and is the first to support multimodal condition joint driven motion generation tasks.