← 返回论文检索
CVPR 2026

RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations

Mochu Xiang, Zhelun Shen, Xuesong Li, Jiahui Ren, Jing Zhang, Chen Zhao, Shanshan Liu, Haocheng Feng, Jingdong Wang, Yuchao Dai

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Humans perceive the 3D world from limited 2D observations. While recent feed-forward generalizable 3D reconstruction models can recover structures from sparse images, they typically represent only observed regions, leaving unseen geometry unmodeled. This raises a fundamental question: Can we infer complete 3D structure from partial observations? We present RnG (Reconstruction and Generation), a feed-forward Transformer that unifies reconstruction and generation by predicting an implicit, complete 3D representation. RnG introduces a reconstruction-guided causal attention mechanism that separates reconstruction and generation at the attention level while treating the KV-cache as an implicit 3D representation. Arbitrary poses can efficiently query this cache to render high-fidelity RGBD outputs. RnG accurately reconstructs visible geometry while generating plausible unseen structures and appearance. It achieves state-of-the-art performance in generalizable 3D reconstruction and novel view generation, while being efficient enough for real-time interactive applications.