Mechanistic Insights into Deferred Semantic Drift in LLMs
Dalian University of Technology
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.57 ↗
摘要
Large Language Models (LLMs) face a fundamental challenge with delayed disambiguation: **How is the meaning of an ambiguous word updated when clarifying context arrives only after it has been processed?** While LLMs possess the latent capacity to resolve such ambiguities—as revealed when a full, non-causal context is provided—their unidirectional architecture prevents immediate updates. We investigate the underlying computational mechanism and show this semantic re-evaluation is deferred to subsequent tokens in a process we term "Deferred Semantic Drift (DSD)". Through targeted analysis of attentional pathways, we find that later tokens actively retrieve context-dependent "informational packets" from the ambiguous word’s value vector to steer the final interpretation. We demonstrate this mechanism in metaphor comprehension and provide causal validation by steering model outputs towards literal or metaphorical meanings via targeted activation interventions. This research uncovers a key computational strategy for meaning construction, offering crucial insights for understanding and guiding the behavior of LLMs.