← 返回论文检索
ACM Multimedia 2025Generative AI: Generative Multimedia

DCNOT: Diffusion-Cascaded Neural Optimal Transport for Scalable Multi-Domain Image-to-Image Translation

Yingzhen Zhang, Jimin Dai, Qianliang Wu, Jian Yang 0003, Lei Luo 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754979 ↗

摘要

Optimal Transport (OT) has emerged as a principled framework for learning mappings between probability distributions by minimizing transportation costs. While neural OT methods have achieved remarkable success in dual-domain (N=2) image-to-image (I2I) translation, their extension to multi-domain settings (N>2) remains challenging due to the quadratic complexity (O(N2)), leading to computational inefficiency and poor scalability. In this work, we propose Diffusion-Cascaded Neural Optimal Transport (DCNOT), a novel approach that reduces the complexity of multi-domain I2I translation to linear (O(N)) leveraging the contracting properties of the forward diffusion process and a cascaded OT strategy. We first prove theoretically that the Wasserstein-2 distance between domains contracts progressively under diffusion noise injection, enabling the alignment of all domains to a shared approximate domain. The remaining distributional shifts are then decomposed into smaller, more tractable gaps bridged via cascaded neural OT mappings, ensuring both efficiency and fidelity. Extensive experiments on synthetic and real-world benchmarks demonstrate that DCNOT achieves state-of-the-art scalability in multi-domain translation while preserving or surpassing the quality of prior OT-based methods. Our work establishes a new paradigm for scalable multi-domain learning with optimal transport.