← 返回论文检索
EMNLP 2025mainmain

Can LLMs Explain Themselves Counterfactually?

Zahra Dehghanighobadi, Asja Fischer, Muhammad Bilal Zafar

Ruhr-Universität Bochum · Ruhr-Universität Bochum and Research Center for Trustworthy Data Science and Security

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.emnlp-main.396 ↗

摘要

Explanations are an important tool for gaining insights into model behavior, calibrating user trust, and ensuring compliance.Past few years have seen a flurry of methods for generating explanations, many of which involve computing model gradients or solving specially designed optimization problems.Owing to the remarkable reasoning abilities of LLMs, *self-explanation*, i.e., prompting the model to explain its outputs has recently emerged as a new paradigm.We study a specific type of self-explanations, *self-generated counterfactual explanations* (SCEs).We test LLMs’ ability to generate SCEs across families, sizes, temperatures, and datasets. We find that LLMs sometimes struggle to generate SCEs. When they do, their prediction often does not agree with their own counterfactual reasoning.