← 返回论文检索
KDD 2025Applied Data Invited Talks

Large Language Models: Architecture and Training. From Next-Word Prediction to Reasoning

Jay Alammar

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3711896.3736805 ↗

摘要

This talk is a highly visual and accessible look at large language models, their architecture, and their training. Attendees will be presented with the intuitions for tens of LLM concepts like tokenizers, the internals of the latest Transformer neural networks, mixture-of-expert models, reward models, reasoning LLMs, model merging, and more.