MuPT: A Generative Symbolic Music Pretrained Transformer
University of Manchester · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Chinese Academy of Sciences · Queen Mary University of London · The Hong Kong University of Science and Technology · University of Macau · Nanjing University · Hong Kong University of Science and Technology · Stanford University · Midea Group (Shanghai) Co.,Ltd. · Montreal Institute for Learning Algorithms, University of Montreal, University of Montreal · 01.AI · Beijing University of Posts and Telecommunications · University of the Chinese Academy of Sciences · Central Conservartory of Music · Peking University · Shanghai Jiaotong University · MBZUAI · Carnegie Mellon University · University of Manchester · Microsoft Research · University of Waterloo
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition.To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks. Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions.