BIMCompNet: Multimodal Dataset for Geometric Deep Learning in Building Information Model
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3758238 ↗
摘要
Building Information Model (BIM) has become a significantly digital platform for representing buildings in the Architecture, Engineering, and Construction (AEC) industry. However, the absence of extensive, class- diverse, and balanced datasets at the BIM component level has limited the development of AI-driven BIM analysis. In this study, BIMCompNet is proposed as a large-scale multimodal dataset from Industry Foundation Classes (IFC), which can learn BIM component geometry features from multiple representation methods, including rendered views, point clouds, mesh structures, voxel grids, and semantic graphs. BIMCompNet is constructed by a standardized two-stage processing pipeline: (1) At the model level, geometry units are normalized to the SI units, models are converted to the IFC format, metadata is anonymized, and components are automatically extracted into individual IFC files. (2) At the component level, semantic labels are corrected, geometry and positioning are aligned, duplicates at model and project levels are removed, and five synchronized modalities (OBJ meshes, multi-view images, point clouds, voxel grids, and heterogeneous IFC graphs) are generated. BIMCompNet comprises 1,304,206 cleaned and labeled components across 87 IFC classes, collected from 1,607 real-world BIM models spanning 14 building types. To mitigate class imbalance, underrepresented classes are merged, and dominant classes are down-sampled to create balanced subsets suitable for robust AI model training and benchmarking. Benchmarking is performed on classification tasks by different models with multiple data modalities. Both the dataset and the processing pipeline will be publicly released to support reproducibility and private dataset extension.