← 返回论文检索
ICML 2026PosterAccept (regular)

MICE-Bench: A Challenging and Comprehensive Benchmark for Multi-Reference Image Creation and Editing

Siqi Luo, Huayu Zheng, Jianghan Shen, Yi Xin, Luxin Xu, Jiyao Liu, Xinyu Zhang, Hang Zhou, Pengyu Xie, Xiaohui Li, Shuo Cao, Yuandong Pu, Junjun He, Bin Fu, Yihao Liu, Yu Qiao, Guangtao Zhai, Yuewen Cao, Xiaohong Liu

Shanghai Jiaotong University · The University of Hong Kong · Nanjing university · Shanghai Artificial Intelligence Laboratory · Fudan University · University of Electronic Science and Technology of China · nanjing university · University of Science and Technology of China · Shanghai AI Laboratory · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Shanghai Aritifcal Intelligence Laboratory · Shanghai Jiao Tong University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

The paradigm of visual generation is rapidly shifting from single-image conditioning toward multi-image conditioning, making the ability to synthesize and edit images based on multiple visual references a critical capability. Despite this trend, existing benchmarks remain largely limited to single-reference scenarios or narrowly defined tasks, leaving model behavior under complex multi-concept composition insufficiently explored. To bridge this gap, we introduce **MICE-Bench**, a comprehensive benchmark for **M**ulti-reference **I**mage **C**reation and **E**diting. The benchmark is designed around three core principles: 1) heterogeneous concept composition across seven visual dimensions; 2) varying levels of constraint density, ranging from dual-concept to seven-concept configurations; 3) concept-centric data construction and benchmark evaluation, enabling fine-grained analysis of interactions among multiple concepts. MICE-Bench consists of 3,119 high-quality test cases within a unified concept space. Using an 8-dimensional evaluation metric, we systematically evaluate 13 state-of-the-art models. Our results show that although closed-source models maintain a clear performance advantage, all models experience notable degradation in concept consistency and physical realism as concept complexity increases.This indicates that current models rely on superficial composition rather than genuine multi-concept synthesis, highlighting substantial room for future improvement.