← 返回论文检索
ICML 2025PosterAccept (poster)

MindCustomer: Multi-Context Image Generation Blended with Brain Signal

Muzhou Yu, Shuyun Lin, Lei Ma, Bo Lei, Kaisheng Ma

Xi'an Jiaotong University · Tsinghua University, Tsinghua University · Peking University · Beijing Academy of Artificial Intelligence · Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Advancements in generative models have promoted text- and image-based multi-context image generation. Brain signals, offering a direct representation of user intent, present new opportunities for image customization. However, it faces challenges in brain interpretation, cross-modal context fusion and retention. In this paper, we present MindCustomer to explore the blending of visual brain signals in multi-context image generation. We first design shared neural data augmentation for stable cross-subject brain embedding by introducing the Image-Brain Translator (IBT) to generate brain responses from visual images. Then, we propose an effective cross-modal information fusion pipeline that mask-freely adapts distinct semantics from image and brain contexts within a diffusion model. It resolves semantic conflicts for context preservation and enables harmonious context integration. During the fusion pipeline, we further utilize the IBT to transfer image context to the brain representation to mitigate the cross-modal disparity. MindCustomer enables cross-subject generation, delivering unified, high-quality, and natural image outputs. Moreover, it exhibits strong generalization for new subjects via few-shot learning, indicating the potential for practical application. As the first work for multi-context blending with brain signal, MindCustomer lays a foundational exploration and inspiration for future brain-controlled generative technologies.