YuE: Scaling Open Foundation Models for Long-Form Music Generation
Hong Kong University of Science and Technology · Beijing University of Posts and Telecommunications · University of Waterloo · Smule, Inc. · Ohio State University · University of the Chinese Academy of Sciences · Mohamed bin Zayed University of Artificial Intelligence · University of Rochester · 01.AI · The Hong Kong University of Science and Technology · Zhejiang University · Queen Mary University of London · 2077AI · University of Manchester · Tencent · Tianjin University · Shanghai Jiao Tong University · Fudan University · JD.com · National University of Singapore · China University of Geoscience Beijing · Wuhan University of Engineering Science · Institute of Artificial Intelligence (TeleAI), China Telecom · Microsoft · ByteDance Inc. · Kuaishou- 快手科技 · CMU, Carnegie Mellon University;ZJU,Zhejiang University · Google DeepMind · Geely Automobile Research Institute (Ningbo) Co., Ltd · University of Manchester · Shanghai Jiaotong University · MBZUAI · Institute of automation, Chinese academy of science, Chinese Academy of Sciences · Department of Electronic Engineering, Tsinghua University · Megvii Inc. · Carnegie Mellon University · Nanjing University · Beihang University · Microsoft Research · Imperial College London
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while maintaining lyrical alignment, coherent musical structure, and engaging vocal melodies with appropriate accompaniment. It achieves this through \textbf{track-decoupled next-token prediction} to overcome dense mixture signals, and \textbf{structural progressive conditioning} for long-context lyrical alignment. In addition, we redesign the \textbf{in-context learning} technique for music generation, enabling bidirectional content creation, style cloning, and improving musicality. Through extensive evaluation, we demonstrate that YuE matches or even surpasses some of the proprietary systems in musicality and vocal agility (as of 2025-01). We strongly encourage readers to \textbf{listen to our demo}\footnote{\url{https://map-yue.github.io/}}.