Noise-Optimized Distribution Distillation for Dataset Condensation
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755602 ↗
摘要
Dataset condensation distills a large dataset into a small synthetic surrogate dataset with similar training efficacy on downstream tasks. Of the existing condensation methods, diffusion-based methods that synthesize surrogate datasets with diffusion models have successfully distilled high-resolution datasets with high training efficacy and satisfactory cross-architectural transferability. However, these methods exhibit a random sampling bias that impairs their performance in dataset condensation settings. We propose a novel dataset condensation method called Noise-Optimized Distribution Distillation (NODD) that mitigates this sampling bias to improve the training performance of synthetic datasets generated with diffusion models. NODD can integrate with existing diffusion-based methods to produce synthetic datasets with enhanced training performance.