← 返回论文检索
ICML 2025PosterAccept (poster)

Elucidating the design space of language models for image generation

Xuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu, JUN WANG, Rong Xiao, Yuan YAO

Hong Kong University of Science and Technology · University of Hong Kong · Intelifusion Inc. · Huawei Noah's Ark Lab · HKUST FYTRI · Intellifusion · HongKong University of Science and Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

The success of large language models (LLMs) in text generation has inspired their application to image generation. However, existing methods either rely on specialized designs with inductive biases or adopt LLMs without fully exploring their potential in vision tasks. In this work, we systematically investigate the design space of LLMs for image generation and demonstrate that LLMs can achieve near state-of-the-art performance without domain-specific designs, simply by making proper choices in tokenization methods, modeling approaches, scan patterns, vocabulary design, and sampling strategies. We further analyze autoregressive models' learning and scaling behavior, revealing how larger models effectively capture more useful information than the smaller ones. Additionally, we explore the inherent differences between text and image modalities, highlighting the potential of LLMs across domains. The exploration provides valuable insights to inspire more effective designs when applying LLMs to other domains. With extensive experiments, our proposed model, **ELM** achieves an FID of 1.54 on 256$\times$256 ImageNet and an FID of 3.29 on 512$\times$512 ImageNet, demonstrating the powerful generative potential of LLMs in vision tasks.

论文信息

会议
ICML 2025
年份
2025
主题
Applications->Computer Vision