← 返回论文检索
ACM Multimedia 2025Experience: Multimedia Applications

From Outline to Detail: An Hierarchical End-to-end Framework for Coherent and Consistent Visual Novel Generation and Assembly

Yilin Zhang 0012, Yanyan Wei, Zhao Zhang 0001, Jicong Fan 0001, Haijun Zhang 0002, Shuicheng Yan

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755541 ↗

摘要

As a form of multimedia creation, visual novel (VN) conveys engaging narratives through the integrated presentation of text, images, and music, and has shown promise across various application domains. Recent advances in generative AI have fueled interest in automating VN creation using LLMs and other foundation models. However, fully end-to-end VN creation (i.e., from user description to executable VN) remains underexplored and presents several key challenges: 1) the hallucination and limited capacity of LLMs hinder the generation of long and coherent plots; 2) current models lack effective mechanisms for ensuring cross-modal consistency between plot, visual, and audio elements. To address these issues, we propose a hierarchical end-to-end framework for automatic VN generation and assembly, which employs an outline-guided autoregressive generation mechanism that transforms high-level user prompts into coherent plots, while a vision LLM-based self-correction mechanism ensures consistency between multimedia assets and plot content. Additionally, we introduce a script validation mechanism to ensure the executable of the final VN application. Experiments demonstrate that our framework generates high-quality VN applications with coherent storylines and consistent multimedia content.