← 返回论文检索
NeurIPS 2025{location} PosterAccept (poster)

Antidistillation Sampling

Yash Savani, Asher Trockman, Zhili Feng, Yixuan Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, Zico Kolter

Carnegie Mellon University · CMU · OpenAI

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Frontier models that generate extended reasoning traces inadvertently produce token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. *Antidistillation sampling* provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's utility.