OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3756871 ↗
摘要
Open-source foundation models are essential for advancing music audio understanding and ensuring access to general-purpose representations for music information retrieval. To this end, we present OMAR-RQ, a model trained with self-supervision via masked token prediction using a large-scale dataset with over 330,000 hours of music audio. We experiment with various input features and quantization options, outperforming existing open self-supervised models in music tagging, pitch estimation, chord recognition, beat tracking, segmentation, and difficulty estimation. Finally, we release our training and evaluation pipelines and model weights at https://github.com/mtg/omar-rq.