Small-to-Large Generalization: Training Data Influences Models Consistently Across Scale
OpenAI · Massachusetts Institute of Technology
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Choice of training data distribution greatly influences model behavior. Yet, inlarge-scale settings, precisely characterizing *how* changes in trainingdata affects predictions is often difficult due to model training costs. Currentpractice is to instead extrapolate from scaled down, inexpensive-to-train proxymodels. However, changes in data do not influence smaller and larger modelsidentically. Therefore, understanding how choice of data affects large-scalemodels raises the question: how does training data distribution influence modelbehavior across compute scale? We find that small- and large-scale languagemodel predictions (generally) *do* highly correlate across choice oftraining data. Equipped with these findings, we characterize how proxy scaleaffects effectiveness in two downstream proxy model applications: dataattribution and dataset selection.
论文信息
- 会议
- ICLR 2025
- 年份
- 2025