Human Motion Generation in 3D Scenes from Open-Ended Textual Instructions with MLLM Planning
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754850 ↗
摘要
Generating human motion in scenes from text aims to synthesize semantically aligned and scene-aware motions. Existing methods have made significant progress by incorporating spatial reasoning and structured generation strategies to connect text descriptions with human-scene interactions. However, they typically rely on simple textual inputs and struggle to comprehend open-ended instructions. There are three key challenges: (1) difficulty in understanding complex instructions due to limited and templated training text annotations; (2) inability to generate natural motions that align with arbitrary trajectories described in text; (3) lack of motion diversity that matches the intended semantics. To address these challenges, we propose PSMo, which consists of two components: the Semantic Planner and the Scene-Aware Motion Generator. The Semantic Planner leverages a Multimodal Large Language Model (MLLM) to parse open-ended instructions, and plans fine-grained motion states aligned with arbitrary trajectories. The scene-aware motion generator adopts the diffusion model with trajectory constraints and a sequential tiling strategy. To enhance motion diversity, we introduce a retrieval-augmented strategy and Scene-Aware Retrieval Attention, which integrates multi-modal features into the generation process. Extensive experiments demonstrate that our method produces high-quality and natural motions under open-ended instructions in scenes.