← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Text-to-Image Generation with Multi-modal Knowledge Graph Construction and Retrieval

Jiawei Meng, Zhengmao Yang, Zhiqiang Liu, Shaokai Chen, Zhizhen Liu, Wen Zhang 0015, Huajun Chen

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754792 ↗

摘要

Current Text-to-Image (T2I) generation methods struggle to accurately create images with complex object relationships and scene compositions. To overcome these challenges, we propose KAIG, a novel text-to-image generative model that integrates a knowledge graph into the image generation process. Unlike traditional models, KAIG uses structured knowledge to enhance the retrieval of relevant information, enabling the generation of high-quality, contextually rich, and semantically consistent images from multi-modal inputs. We introduce a two-stage training strategy: first, condition adapters are trained to align multi-modal inputs, followed by fine-tuning the entire diffusion model. This approach ensures precise alignment between retrieved conditions and the image generation process, leading to an efficient and scalable pipeline. Our experiments on two popular datasets, MS-COCO and CUB-200-2011, show that KAIG consistently outperforms existing methods in both image quality and consistency. Notably, KAIG can seamlessly integrate with any pre-trained diffusion model, requiring minimal additional training while achieving superior results. Ultimately, KAIG demonstrates strong potential for addressing key limitations in current T2I models and advancing the field of image synthesis.