← 返回论文检索
ACM Multimedia 2025Generative AI: Generative Multimedia

S2-Edit3DV: Diffusion-Guided Style Meets Structure for Consistent Multi-View 3D Video Generation

Yuqi Chen, Xiubo Liang, Yu Zhao, Hongzhi Wang 0009, Weidong Geng

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755796 ↗

摘要

Consistently stylizing and editing 3D objects from multiple viewpoints is crucial for immersive applications such as virtual reality, augmented reality, and digital entertainment. Nevertheless, existing methods frequently face significant challenges, including inconsistent textures, pronounced drifting artifacts, and compromised geometric integrity when rendered from various perspectives. To effectively address these limitations, we introduce S2-Edit3DV, a novel diffusion-guided framework that reframes multi-view 3D objects editing as a temporally coherent video editing problem. By exploiting the robust single-view generative capabilities of SV3D, our approach reliably propagates initial style edits across different viewpoints, substantially mitigating drifting artifacts prevalent in current video-based editing methods. To further enhance semantic precision and structural preservation, we propose two innovative techniques: Attention-based Differential Style Injection (ADSI) and Adaptive Structural-aware Plug-and-Play (AS-PnP). ADSI utilizes attention-driven semantic embeddings for adaptive and precise style injection, effectively reducing semantic hallucinations. AS-PnP strategically modulates stylized latent features, balancing artistic expression with strict structural coherence. Comprehensive evaluations and ablation studies demonstrate that our proposed framework significantly enhances multi-view consistency, preserves fine-grained geometric details, and ensures accurate semantic alignment, showcasing superior performance and practical value for generating high-quality, creatively stylized, and structurally robust objects.