← 返回论文检索
ACM Multimedia 2025Experience: Interactions and Quality of Experience

A Multimodal Evaluation Framework for Spatial Audio Playback Systems: From Localization to Listener Preference

Changhao Pan, Wenxiang Guo, Yu Zhang 0126, Zhiyuan Zhu, Zhetao Chen, Han Wang 0019, Zhou Zhao 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755571 ↗

摘要

Spatial audio playback defines immersive listening. However, objective evaluation methods for perceptual dimensions like sound field and sound image remain underdeveloped, hindered by the lack of fine-grained spatial audio datasets and the neglect of echoes and reverberation in diverse playback conditions. To address these challenges, we propose MESA, a multi-modal evaluation framework for spatial audio systems, and introduce PSA-MOS, a high-quality multi-scene spatial audio dataset. Specifically: 1) PSA-MOS provides 50 hours of high-quality spatial audio recordings spanning 6 playback scenarios and 7 device types, with detailed localization annotations and fine-grained MOS ratings across four perceptual dimensions. 2) We develop SAE-Encoder, a spatial audio encoder that captures both acoustic-spatial cues and fine-grained perceptual patterns. 3) MESA integrates visual scene context to enhance evaluation robustness through echo and reverberation modeling. Experimental results demonstrate that SAE-Encoder achieves superior performance in SELD tasks. With a two-stage training strategy, MESA exhibits strong correlation with human perceptual assessments, effectively guiding spatial audio quality optimization. The demos are available at https://david-pigeon.github.io/mesaDemo.