← 返回论文检索
NeurIPS 2025{location} Oral PosterAccept (Oral)

Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research

A. Feder Cooper, Christopher Choquette-Choo, Miranda Bogen, Kevin Klyman, Matthew Jagielski, Katja Filippova, Ken Liu, Alex Chouldechova, Jamie Hayes, Yangsibo Huang, Eleni Triantafillou, Peter Kairouz, Nicole Mitchell, Niloofar Mireshghallah, Abigail Jacobs, James Grimmelmann, Vitaly Shmatikov, Christopher De Sa, I Shumailov, Andreas Terzis, Solon Barocas, Jennifer Wortman Vaughan, danah boyd, Yejin Choi, Sanmi Koyejo, Fernando Delgado, Percy Liang, Daniel Ho, Pamela Samuelson, Miles Brundage, David Bau, Seth Neel, Hanna Wallach, Amy Cyphert, Mark Lemley, Nicolas Papernot, Katherine Lee

Stanford University · OpenAI · Center for Democracy & Technology · Anthropic · Research, Google · Microsoft · Google DeepMind · Google · Google Research · UCSD · University of Michigan · Cornell University · University of Toronto · Microsoft Research; Cornell University · Microsoft Research · Data & Society Research Institute · UW => Stanford / NVIDIA · Stanford University / Virtue AI · Stanford Law · University of California, Berkeley · Northeastern University · West Virginia University · University of Toronto and Vector Institute

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific information from a generative-AI model's parameters, e.g., a particular individual's personal data or the inclusion of copyrighted content in the model's training data. Unlearning is also proposed as a way to prevent a model from generating targeted types of information in its outputs, e.g., generations that closely resemble a particular individual's data or reflect the concept of "Spiderman." Both of these goals--the targeted removal of information from a model and the targeted suppression of information from a model's outputs--present various technical and substantive challenges. We provide a framework for ML researchers and policymakers to think rigorously about these challenges, identifying several mismatches between the goals of unlearning and feasible implementations. These mismatches explain why unlearning is not a general-purpose solution for circumscribing generative-AI model behavior in service of broader positive impact.