AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models
Tsinghua University · Beijing Normal–Hong Kong Baptist University · Waseda University · Westlake University · Huazhong University of Science and Technology · Beijing Jiaotong University · The Hong Kong Polytechnic University · Qinghai University · University of Rochester · New York University · Hong Kong Polytechnic University · Beihang University · Shanghai Jiao Tong University · NUS · Institute of automation, Chinese academy of science · University of Science and Technology of China · Nanyang Technological University · The Hong Kong University of Science and Technology · ByteDance Inc. · Massachusetts Institute of Technology · Head of AI @ Squirrel Ai Learning · Chinese University of Hong Kong, Shenzhen · Zhejiang University · Indiana University
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
The rapid development and widespread adoption of Audio Large Language Models (ALLMs) require a rigorous assessment of their trustworthiness. However, existing evaluation frameworks, primarily designed for text, are not equipped to handle the unique vulnerabilities introduced by audio’s acoustic properties. We find that significant trustworthiness risks in ALLMs arise from non-semantic acoustic cues, such as timbre, accent, and background noise, which can be used to manipulate model behavior. To address this gap, we propose AudioTrust, the first framework for large-scale and systematic evaluation of ALLM trustworthiness concerning these audio-specific risks. AudioTrust spans six key dimensions: fairness, hallucination, safety, privacy, robustness, and authenticition. It is implemented through 26 distinct sub-tasks and a curated dataset of over 4,420 audio samples collected from real-world scenarios (e.g., daily conversations, emergency calls, and voice assistant interactions), purposefully constructed to probe the trustworthiness of ALLMs across multiple dimensions. Our comprehensive evaluation includes 18 distinct experimental configurations and employs human-validated automated pipelines to objectively and scalably quantify model outputs. Experimental results reveal the boundaries and limitations of 14 state-of-the-art (SOTA) open-source and closed-source ALLMs when confronted with diverse high-risk audio scenarios, thereby offering critical insights into the secure and trustworthy deployment of future audio models. Our platform and benchmark are publicly available at https://github.com/JusperLee/AudioTrust.