TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models
University of Notre Dame · Mohamed bin Zayed University of Artificial Intelligence · Emory University · Sichuan University · Huazhong University of Science and Technology · University of Illinois at Urbana-Champaign · University of Cambridge · Department of Computer Science, University of Maryland, College Park · University of California, Los Angeles · University of Washington · CMU, Carnegie Mellon University · National University of Singapore · Salesforce Research · Department of Computer Science, UT Austin · University of Waterloo · University of Queensland · CMU · UNC-Chapel Hill · Department of Computer Science, University of Washington · NTU Singapore · Department of Computer Science · University of Southern California · University of Washington, Seattle · Massachusetts General Hospital · Massachusetts Institute of Technology · School of Computer Science, Carnegie Mellon University · KAUST · Texas A&M University - College Station · Arizona State University · MBZUAI · Lehigh University · University of Maryland · University of Miami · IBM Research · Stanford University · Allen Institute for AI · Ohio State University · Johns Hopkins University/NVIDIA · UNC Chapel Hill · Simon Fraser University · Microsoft Research · CISPA Helmholtz Center for Information Security · University of Illinois, Chicago · IBM Research AI · University of Illinois, Urbana Champaign · Berkeley
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Generative foundation models (GenFMs), such as large language models and text-to-image systems, have demonstrated remarkable capabilities in various downstream applications. As they are increasingly deployed in high-stakes applications, assessing their trustworthiness has become both a critical necessity and a substantial challenge. Existing evaluation efforts are fragmented, rapidly outdated, and often lack extensibility across modalities. This raises a fundamental question: how can we systematically, reliably, and continuously assess the trustworthiness of rapidly advancing GenFMs across diverse modalities and use cases? To address these gaps, we introduce TrustGen, a dynamic and modular benchmarking system designed to systematically evaluate the trustworthiness of GenFMs across text-to-image, large language, and vision-language modalities. TrustGen standardizes trust evaluation through a unified taxonomy of over 25 fine-grained dimensions—including truthfulness, safety, fairness, robustness, privacy, and machine ethics—while supporting dynamic data generation and adaptive evaluation through three core modules: Metadata Curator, Test Case Builder, and Contextual Variator. Taking TrustGen into action to evaluate the trustworthiness of 39 models reveals four key insights. (1) State-of-the-art GenFMs achieve promising overall trust performance, yet significant limitations remain in specific dimensions such as hallucination resistance, fairness, and privacy preservation. (2) Contrary to prevailing assumptions, open-source models now rival and occasionally surpass proprietary systems in trustworthiness metrics. (3) The trust gap among top-performing models is narrowing, likely due to increased industry convergence on best practices. (4) Trustworthiness is not an isolated property; it interacts complexly with other behaviors, such as helpfulness and ethical decision-making. TrustGen is a transformative step toward standardized, scalable, and actionable trustworthiness evaluation, supporting dynamic assessments across diverse modalities and trust dimensions that evolve alongside the generative AI landscape.