← 返回论文检索
ICLR 2026PosterAccept (Poster)

W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image Editing

Jiahui Sun, Weining Wang, Mingzhen Sun, Peiyao Wang, Xinxin Zhu, Jing Liu

Institute of automation, Chinese academy of science, Chinese Academy of Sciences · Institute of automation, Chinese Academy of Sciences · ByteDance Inc. · Beijing Institute of Technology · , Institute of automation, Chinese academy of science · Institute of automation, Chinese academy of science

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

While recent advances in Diffusion Transformers (DiTs) have significantly advanced text-to-image generation, text-driven image editing remains challenging. Existing approaches either struggle to balance structural preservation with flexible modifications or require costly fine-tuning of large models. To address this, We introduce W-Edit, a training-free framework for text-driven image editing based on wavelet-based frequency-aware feature decomposition. W-Edit employs wavelet transforms to decompose diffusion features into multi-scale frequency bands, disentangling structural anchors from editable details. A lightweight replacement module selectively injects these components into pretrained models, while an inversion-based frequency modulation strategy refines sampling trajectories using structural cues from attention features. Extensive experiments demonstrate that W-Edit achieves high-quality results across a wide range of editing scenarios, outperforming previous training-free approaches. Our method establishes frequency-based modulation as both a sound and efficient solution for controllable image editing.