← 返回论文检索
ACM Multimedia 2025Datasets

T23D-QA: An Open Dataset and Benchmark for Text-driven 3D Generation Quality Assessment

Haohui Li, Bowen Qu, Wei Gao 0003

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3758302 ↗

摘要

Emerging text-to-3D generative models based on large diffusion backbones have markedly lowered the barrier to high-quality 3D asset creation, yet rigorous quantitative evaluation remains elusive. We introduce T23D-QA, a novel open benchmark that couples diverse prompts, multiple generation paradigms, and fine-grained human judgements for text-conditioned 3D synthesis. The dataset comprises 1,710 textured meshes produced by nine state-of-the-art pipelines spanning feed-forward, optimization-based, and view-reconstruction families. Assets are driven by a tri-categorical prompt suite: single-object, multi-object, and primitive-anchored, covering 160 ShapeNet-level object classes. Each mesh is rated by 20 participants along three orthogonal dimensions: geometry, texture, and alignment. Building upon this corpus, we propose an evaluator that decouples multi-modal features via cross-attention. On T23D-QA, our baseline surpasses the strongest published metric by 8.1% (geometry), 6.1% (texture), and 1.9% (alignment) in Spearman rank correlation. Dataset and code are publicly available at https://t23d-qa.github.io to foster reproducible research.