← 返回论文检索
ACM Multimedia 2025Demos and Videos

An Aesthetic Cultural Relic Poster Generation Framework Based on Multi-target Learning and Multimodal Large Language Model

Mohan Zhang, Qianqian Hu, Chuhan Li, Yanxiu Dan, Shenglan Cui, Fang Liu 0002

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754461 ↗

摘要

This paper presents CrePoster, a data-driven framework to generate aesthetic posters for Chinese cultural relics, aiming to enhance the exhibition experience and promote cultural spread. CrePoster comprises three modules: (1) object segmentation module, (2) content generation module, and (3) poster generation module. Upon processing a cultural relic image, the object segmentation module first leverages a cascaded U2Net-SAM structure to obtain the visual target. Secondly, the content generation module utilizes a multi-target learning-enabled caption generator to produce professional captions. Thirdly, the Multimodal Large Language Model (MLLM) based poster generation module adaptively creates aesthetic parameters, including layout and color scheme, ultimately rendering them into refined posters.