← 返回论文检索
ICML 2024Spotlight PosterAccept (Spotlight)

Variational Learning is Effective for Large Deep Networks

Yuesong Shen, Nico Daheim, Bai Cong, Peter Nickl, Gian Maria Marconi, Bazan Raoul, Rio Yokota, Iryna Gurevych, Daniel Cremers, Khan Emtiyaz, Thomas Moellenhoff

Technical University Munich · Technische Universität Darmstadt · Tokyo Institute of Technology · RIKEN Center for AI Project · RIKEN · Tokyo Institute of Technology, Tokyo Institute of Technology · Technical University of Darmstadt / Mohamed bin Zayed University of Artificial Intelligence · TU Munich

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

We give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks. We show that an optimizer called Improved Variational Online Newton (IVON) consistently matches or outperforms Adam for training large networks such as GPT-2 and ResNets from scratch. IVON's computational costs are nearly identical to Adam but its predictive uncertainty is better. We show several new use cases of IVON where we improve finetuning and model merging in Large Language Models, accurately predict generalization error, and faithfully estimate sensitivity to data. We find overwhelming evidence that variational learning is effective. Code is available at https://github.com/team-approx-bayes/ivon.