← 返回论文检索
ACL 2024aclfindings

When is a Language Process a Language Model?

Li Du, Holden Lee, Jason Eisner, Ryan Cotterell

Johns Hopkins University · Microsoft and Johns Hopkins University · Swiss Federal Institute of Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2024.findings-acl.659 ↗

摘要

A language model may be viewed as a \Sigma-valued stochastic process for some alphabet \Sigma.However, in some pathological situations, such a stochastic process may “leak” probability mass onto the set of infinite strings and hence is not equivalent to the conventional view of a language model as a distribution over ordinary (finite) strings.Such ill-behaved language processes are referred to as *non-tight* in the literature.In this work, we study conditions of tightness through the lens of stochastic processes.In particular, by regarding the symbol as marking a stopping time and using results from martingale theory, we give characterizations of tightness that generalize our previous work [(Du et al. 2023)](https://arxiv.org/abs/2212.10502).