Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
National Taiwan University · ByteDance · University of Texas at Austin · CMU, Carnegie Mellon University · Carnegie Mellon University · Nanyang Technological University · School of Computer Science, Carnegie Mellon University · TTI-Chicago · Institut National de la Recherche Scientifique (INRS - EMT) · NVIDIA · Department of computer science and informational engineering, National Taiwan University · CancerFree · National Taiwan University, Genibuilder · Toyota Technological Institute at Chicago · ASAPP · Renmin University of China · NVIDIA Research · University of Texas, Austin
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication gaps and facilitating more intuitive interactions. However, the absence of a comprehensive evaluation benchmark poses a significant challenge. We present Dynamic-SUPERB Phase-2, an open and evolving benchmark for the comprehensive evaluation of instruction-based universal speech models. Building upon the first generation, this second version incorporates 125 new tasks contributed collaboratively by the global research community, expanding the benchmark to a total of 180 tasks, making it the largest benchmark for speech and audio evaluation. While the first generation of Dynamic-SUPERB was limited to classification tasks, Dynamic-SUPERB Phase-2 broadens its evaluation capabilities by introducing a wide array of novel and diverse tasks, including regression and sequence generation, across speech, music, and environmental audio. Evaluation results show that no model performed well universally. SALMONN-13B excelled in English ASR and Qwen2-Audio-7B-Instruct showed high accuracy in emotion recognition, but current models still require further innovations to handle a broader range of tasks. We open-source all task data and the evaluation pipeline at https://github.com/dynamic-superb/dynamic-superb.