← 返回论文检索
ACM Multimedia 2025Open Source Software

OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction

Pablo Alonso-Jiménez, Pedro Ramoneda, Recep Oguz Araz, Andrea Poltronieri, Dmitry Bogdanov

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3756871 ↗

摘要

Open-source foundation models are essential for advancing music audio understanding and ensuring access to general-purpose representations for music information retrieval. To this end, we present OMAR-RQ, a model trained with self-supervision via masked token prediction using a large-scale dataset with over 330,000 hours of music audio. We experiment with various input features and quantization options, outperforming existing open self-supervised models in music tagging, pitch estimation, chord recognition, beat tracking, segmentation, and difficulty estimation. Finally, we release our training and evaluation pipelines and model weights at https://github.com/mtg/omar-rq.