← 返回论文检索
ACM Multimedia 2025Generative AI: Generative Multimedia

MPPR: Memory-Prior-based Prompt Refinement in Continuous Space for Advanced Text-to-Image Generation

Zhibing Zhang, Jiantao Lin, Cangqi Zhou, Rui Xia

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755280 ↗

摘要

Refining user-provided natural language prompts allows users to more easily obtain their desired outputs in text-to-image generation. Existing automatic prompt refinement methods predominantly take discrete, human-engineered high-quality prompts as the final optimization target. However, human-engineered prompts are based on human intuition and derived through limited interaction with generative models, which fails to bridge the gap between human preferences and model preferences. Additionally, this discrete optimization target limits the information capacity of the conditional inputs fed to the generative model, leading to suboptimal outcomes. Therefore, we propose an end-to-end prompt optimization method that interacts directly with generative models, eliminating the need for human involvement. The optimization process takes high-quality images as the target and uses the internal states of generative models as optimization signals. This optimizes prompts in a way that aligns more naturally with the model's generation process and producing continuous representations as the final refined prompt. We also introduce a memory module to store common features of high-quality prompts as prior knowledge to guide optimization in continuous space, enabling it to be more efficient. This memory-p rior-based p rompt r efinement in continuous space (MPPR) not only bridges the gap between human preferences and model preferences, but also resolves the issue of insufficient information in the inputs provided to the generative model. Extensive experiments show that our method achieves better performance compared to the state-of-the-art baselines.