FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations
King Abdullah University of Science and Technology (KAUST) · Hamad Bin Khalifa University (HBKU) · King Abdullah University of Science and Technology · Miami University (OH) · KAUST
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large-language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes (e.g., kitchens, living rooms, bedrooms, bathrooms, and others), encoded symbolically in JSON or XML layouts. The benchmark covers core spatial tasks, including distance measurement, visibility, path finding, and object placement within constrained spaces. Our results across a variety of frontier open-source and commercial LLMs reveal that while models may succeed on shallow queries, they often fail to respect physical constraints and preserve spatial coherence, though they remain mostly robust to small spatial perturbations. FloorplanQA uncovers a blind spot in today’s LLMs: inconsistent reasoning about indoor layouts. We hope this benchmark inspires new work on language models that can accurately infer and manipulate spatial and geometric properties in practical settings.