← 返回论文检索
ICML 2026PosterAccept (regular)

Context-level Language Modeling by Learning Predictive Context Embeddings

beiya dai, Yuliang Liu, Daozheng Xue, Yunchong Song, Qipeng Guo, Kai Chen, Xinbing Wang, Bowen Zhou, Zhouhan Lin

Shanghai Jiao Tong University · Nanjing University · Zhiyuan College, Shanghai Jiao Tong University, Shanghai Jiaotong University · Shanghai Artificial Intelligence Laboratory · Fudan University · Shanghai AI Laboratory · Tsinghua University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

We propose ContextLM, a framework that implicitly learns multi-token prediction by augmenting standard pretraining with an intrinsic next-context prediction objective. ContextLM builds a language model on top of context embeddings that span multiple tokens, enabling better next-token prediction by predicting the next context. Our model is fully compatible with standard autoregressive, token-by-token evaluation paradigms (e.g., perplexity). Extensive experiments with GPT-2 and Pythia backbones (up to 1.5B parameters and 300B training tokens) reveal that ContextLM shifts the Pareto frontier of scaling laws, exhibiting superior efficiency in parameters, training tokens, and FLOPs. Our results show that ContextLM could already achieve the baseline perplexity using 39\% fewer parameters and demonstrates robust generalization improvements on extensive downstream tasks under equivalent parameter counts.