DEFT: Demystifying VLN Failures via a Unified Dual-View Explainability Framework for LLM-based Agents
University of the Chinese Academy of Sciences · Institute of Software, Chinese Academy of Sciences
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.1363 ↗
摘要
Large Language Models (LLMs) have emerged as central planners in Vision-and-Language Navigation (VLN), yet their complexity increasingly obscures their internal decision-making. Existing interpretability methods typically isolate temporal criticality from feature salience, creating an alignment gap and failing to account for the behavioral instability of black-box agents. To address this, we propose DEFT, a unified dual-view framework that demystifies agent behavior by jointly analyzing \textit{when} a decision is pivotal and \textit{what} visual evidence grounds it. Featuring a dual-head architecture with a shared latent representation, DEFT employs a \textit{Mask Head} for counterfactual-based criticality detection and an \textit{Action Head} that leverages an ensemble of surrogates to recover robust visual cues. Extensive experiments on MatterPort3D across three LLM-based agents demonstrate that DEFT outperforms baselines in both temporal and feature fidelity. User studies further validate its utility, showing 78% alignment with human intuition.