← 返回论文检索
NeurIPS 2025{location} PosterAccept (poster)

Abstract Counterfactuals for Language Model Agents

Edoardo Pona, Milad Kazemi Mehrabadi, Yali Du, David Watson, Nicola Paoletti

King's College London, University of London · King’s College London · King‘s College London · King's College London · Royal Holloway, University of London

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Counterfactual inference is a powerful tool for analysing and evaluating autonomous agents, but its application to language model (LM) agents remains challenging. Existing work on counterfactuals in LMs has primarily focused on token-level counterfactuals, which are often inadequate for LM agents due to their open-ended action spaces. Unlike traditional agents with fixed, clearly defined action spaces, the actions of LM agents are often implicit in the strings they output, making their action spaces difficult to define and interpret. Furthermore, the meanings of individual tokens can shift depending on the context, adding complexity to token-level reasoning and sometimes leading to biased or meaningless counterfactuals. We introduce \emph{Abstract Counterfactuals}, a framework that emphasises high-level characteristics of actions and interactions within an environment, enabling counterfactual reasoning tailored to user-relevant features. Our experiments demonstrate that the approach produces consistent and meaningful counterfactuals while minimising the undesired side effects of token-level methods. We conduct experiments on text-based games and counterfactual text generation, while considering both token-level and latent-space interventions.