FreeCAD: A Multimodal Framework for 3D CAD Model Generation from Free-Form Prompts
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755550 ↗
摘要
Developing computer-aided design (CAD) generation models has significantly enhanced design efficiency, facilitating innovation and transformation in the design industry. Existing methods typically require users to input prompts in a specific format, such as text descriptions or images, limiting their broader application in diverse scenarios. To address this limitation, we introduce FreeCAD, a user-friendly CAD generation framework that supports free-form inputs, including text descriptions and/or images, enabling users to express their design intentions more flexibly. Specifically, we propose a Large Language Models (LLMs)-based Text Translator, which effectively increases the success rate of generating CAD models by converting users' diversified requests for the same object into a unified expression. Additionally, the Multi-View Representation Fusion (MVRF) module enables the network to capture richer interaction information across views, facilitating the generation of more fine-grained CAD models. To support the training of FreeCAD, we construct a multimodal dataset RealCAD, comprising text, image, and CAD triplets, where the images are derived from the 3D printed products of CAD models. Extensive experiments demonstrate that FreeCAD consistently outperforms the existing state-of-the-art (SOTA) methods in multiple tasks.