← 返回论文检索
ICML 2026PosterAccept (regular)

Temper-Then-Tilt: Principled Unlearning for Generative Models through Tempering and Classifier Guidance

Jacob L. Block, Mehryar Mohri, Aryan Mokhtari, Sanjay Shakkottai

The University of Texas at Austin · Google Research and Courant Institute of Mathematical Sciences · UT Austin · University of Texas at Austin

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

We study machine unlearning in large generative models by framing the task as density ratio estimation to a target distribution rather than supervised fine-tuning. While classifier guidance is a standard approach for approximating this ratio and can succeed in general, we show it can fail to faithfully unlearn with finite samples when the forget set represents a sharp, concentrated data distribution. To address this, we introduce **Temper-Then-Tilt Unlearning (T3-Unlearning)**, which freezes the base model and applies a two-step inference procedure: (i) *tempering* the base distribution to flatten high-confidence spikes, and (ii) *tilting* the tempered distribution using a lightweight classifier trained to distinguish retain from forget samples. Our theoretical analysis provides finite-sample guarantees linking the surrogate classifier's risk to unlearning quality, proving that tempering is necessary to successfully unlearn for concentrated distributions. Empirical evaluations on the TOFU benchmark demonstrate that T3-Unlearning improves forget quality and generative utility over existing baselines, while training only a fraction of the parameters with a minimal runtime.