We propose Gaussian point splatting, a stochastic method to render Gaussian splats that scales extremely well to scenes with many Gaussians. Our core idea is to sample pixel-sized, opaque points from the Gaussians and to splat them to a framebuffer using 64–bit atomics. Through parallel programming primitives, we achieve an even distribution of the workload across millions of threads. Since these threads splat points independently, multiple points may splat to the same pixel. That makes it non-trivial to determine how many points should be splatted for a Gaussian or how they should be distributed to achieve the desired opacity. We successfully formalize and solve these problems, thus keeping our renders faithful to the original Gaussian splatting. To further accelerate our method, we employ hierarchical frustum and occlusion culling. Our method renders hundreds of millions of Gaussians in real time. The only differences compared to the original Gaussian splatting are slight noise and differences in aliasing.
论文检索
输入标题、作者或关键词,从 326 篇学术成果中精准定位
We present a framework for uncertainty-aware geometry processing on Gaussian Process Implicit Surfaces (GPIS), enabling computations directly on such probabilistic representations of shapes. In contrast to classical geometry processing pipelines that assume deterministic surface meshes or point clouds, our approach considers uncertainty in the input data and defines analogs of fundamental differential operators-gradient, divergence, and Laplacian- that account for the distribution of plausible geometries encoded by the GPIS. Leveraging the Kac-Rice formula, we embed computations from random surfaces into a volumetric Cartesian domain, enabling efficient evaluation of expected integrals and differential operators. The proposed approach bridges classical surface PDE-based geometry processing and volumetric representations, enabling a principled handling of noise and ambiguity for various downstream geometry processing tasks.
Generalized winding numbers provide a robust measure of point insidedness for 3D surfaces—whether open, self-intersecting, or non-manifold—and are central to numerous geometry processing tasks. However, existing methods trade off between accuracy and computational efficiency, limiting their use in interactive and large-scale applications. We introduce a new formulation and algorithm for computing generalized winding numbers that is both fast and accurate to arbitrary precision, applicable to meshes and parametric surfaces. Our approach expresses the winding number as the sum of two intuitive geometric quantities: the signed number of ray-surface intersections and a boundary integral over the surface's projection onto the unit sphere. This insight leads to an efficient discretization that avoids expensive surface integrals and spherical arrangements. For meshes, our method achieves average speedups of 22X on a CPU compared to the fastest precise methods and 3X compared to the fastest approximation method, while maintaining full precision. On a GPU, for moderately complex meshes we reach a throughput of 109 queries per second, or 4K generalized winding number slices at 120 FPS (13X faster than a naïve GPU method). For parametric surfaces, our method is on average 5.6X faster than the state-of-the-art method, with the same precision. Our method naturally handles complex topologies and non-manifold inputs. We extensively validate its accuracy, robustness, and time performance. Our code is available at https://github.com/MartensCedric/antipodal.
The generalized winding number (GWN) is a scalar field that supports robust containment queries on curved geometry, including non-watertight, overlapping, and nested boundary representations. While queries can be easily parallelized over samples, direct evaluation on parametric curves and surfaces remains costly for large and complex models. Fast, state-of-the-art GWN approaches leverage a spatial index to approximate the GWN, typically coupled with a Taylor expansion which approximates the GWN contribution for far clusters of geometric primitives. However, such methods operate only on discrete inputs such as triangle meshes and point clouds, and would introduce containment errors near boundaries if applied to curved input. We extend support for fast GWN evaluation over arbitrary collections of NURBS curves in 2D and trimmed NURBS patches in 3D via a Bounding Volume Hierarchy that stores efficiently precomputed moment data in the hierarchy nodes. When querying the hierarchy, approximations for far clusters are used alongside direct evaluation for nearby NURBS primitives, achieving sub-linear complexity while preserving the geometric features in the vicinity of the query point. Central to our performance improvements is an adaptive subdivision strategy for NURBS primitives during a preprocessing phase, creating better spatial partitions while retaining the same accuracy for containment decisions as a direct evaluation. We demonstrate the performance and accuracy of our approach across a large collection of 2D and 3D datasets.
We revisit the computation of 3D generalized winding numbers, a useful measure for inside-outside classification on triangle meshes with gaps, self-intersections, and open boundaries. At the core of our new method is an analytical reduction of the surface integral that defines the winding number, resulting in a single ray-mesh intersection test and an elementary sum over boundary edges per evaluation. This construction is orders of magnitude more efficient than the state of the art in practice, which we show in an extensive performance benchmark. Conveniently, the method also reduces to the best-available asymptotic complexity in the worst case, and it introduces no approximations apart from floating-point errors. Our algorithm is conceptually simple to understand, straightforward to implement and debug, and it works reliably even on extremely noisy and corrupt input geometry.
Shape-changing displays typically lose pixel density as surface area expands, limiting their usability. We introduce MorphSkein, a shape-changing after-image display that preserves initial density (1.44 px/cm2) across naturally occurring axisymmetric shapes generated by spinning cables (troposkeins). The system uses a telescopic pole and four LED-strip rewinders on a rotating base. As the strips spin, centrifugal force forms troposkeins, and persistence of vision creates 360°-visible displays, while adjusting pole height and strip lengths changes their shape. As surface area grows, pixel density is preserved vertically by releasing new rows from the rewinders and horizontally by rendering extra columns per revolution with the strips. This keeps comparable density along the central horizontal line of the display, with naturally higher density toward the top and bottom where the troposkein curves inward. Because the technique relies on a mathematical model assuming ideal troposkein geometry, angular velocity becomes critical: incorrect speeds degrade pixel density accuracy, axisymmetric shape fidelity, or both. Interpolation of experimental data shows that 69.44% of reachable troposkein configurations achieve ⩾90% density accuracy and shape fidelity for at least one operating speed. Remaining cases degrade due to insufficient motor speed or limited MCU speed and LED refresh rate. Limitations and improvements are discussed.
Holography offers unique advantages for delivering perceptual realism while preserving compact form factors in VR/AR. Its perceptual quality, however, hinges on encoding rich wavefronts of photorealistic scenes into interference patterns and then incoherently multiplexing the resulting wave fields for perception. Existing CGH paradigms decouple radiance estimation from wave propagation by pre-rendering radiance on discretized scene sectors. This separation between radiometric and wave-optical computation inherently limits the range of focus cues and visual effects that can be faithfully reproduced, including depth- and view-continuity, and physically based material behaviors such as glossy or mirror-like reflection and refraction. We present a physically accurate yet computationally efficient wave optics rendering framework leveraging path tracing to encode full 3D visual cues into phase holograms. Specifically, we employ a Monte Carlo method to solve both the rendering equation and the Rayleigh-Sommerfeld integral simultaneously. Our algorithm is fully compatible with modern graphics techniques and can generate multiple time-multiplexed random holograms with minimal additional time cost via Path Reuse. By employing a fast approximation with an ambient radiance cache, we realize an order of magnitude convergence speed improvement. The resulting coherent wave fields that inherently encode comprehensive visual effects are converted into phase-only holograms under complex-amplitude supervision. Through extensive simulations and experimental validations on a spatial light modulator-based display prototype, we demonstrate faithful holographic reconstructions of natural 3D cues and complex materials, including realistic defocus blur, view-dependent effects, as well as appearance highlights and reflections.
Volumetric additive manufacturing promises near-instantaneous fabrication of 3D objects, yet achieving high fidelity at the micro-scale remains challenging due to the complex interplay between optical diffraction and chemical effects. We present Single-View Holographic Volumetric Additive Manufacturing (SHVAM), a mechanically static system that shapes volumetric dose distributions using time-multiplexed, phase-only holograms projected from a single optical axis. To achieve high resolution with SHVAM, we formulate hologram synthesis as a coupled inverse problem, integrating a differentiable wave-optical forward model with a simplified photochemical model that explicitly captures inhibitor diffusion and non-linear dose response. Optimizing hologram sequences under these coupled constraints allows us to pre-compensate for chemical blur, yielding higher print fidelity than optical-only optimization. We demonstrate the efficacy of SHVAM by fabricating simple 2D and 3D structures with lateral feature sizes of approximately 10 μm within a 0.8 mm × 0.8 mm × 3 mm volume in seconds.
Optically recorded analog holograms can reconstruct photorealistic three-dimensional (3D) images without the need for specialized eyewear. Computer-generated holograms (CGHs) are created by simulating the holographic recording process digitally rather than capturing them optically. Large-scale 3D still-image reconstruction with wide-viewing-zone can be achieved by mapping the amplitude or phase profiles of CGHs onto diffractive optical elements (DOEs), which modulate in-plane wavefront distribution using wavelength-scale surface structures. However, DOEs exhibit limited wavelength selectivity, interacting with a broad spectral range beyond their design wavelength. This results in reduced overall transmittance and typically requires three separate CGHs for full-color reconstruction. In this work, we present an "invisible holographic window," a transparent, surface-relief CGH patterned directly on glass via laser grayscale lithography. By encoding the real component of the interference pattern and reducing the phase-modulation range, the surface-relief structure transitions from deeply wrapped, jagged phase profiles to shallower and smoother sinusoidal phase profiles, which is associated with improved optical transmittance. Furthermore, a spatial-frequency-domain band-division multiplexing strategy is applied to support crosstalk-free, full-color 3D image reconstruction from a single transparent CGH, albeit with a viewing angle reduction. This platform generates photorealistic full-color 3D still-images in real space, offering new possibilities for transparent augmented-reality display interfaces such as those integrated into storefront windows, office glass partitions, and museum display cases. Our findings pave the way for advanced holographic display technologies and promise to accelerate research in the field.
For large-scale modal analysis problems, Component Mode Synthesis (CMS) methods are very attractive, as they reduce the global problem into smaller subproblems on substructures. However, the substructure bases do not span the desired solution space efficiently, thus the error decreases slowly as the number of substructure eigenmodes increases. We demonstrate that a much more effective subspace can be constructed by combining substructure eigenmodes from multiple spatially staggered partitions of the input domain. To further accelerate the method, we replace the bases on the substructure interfaces by low-frequency approximations excited from the substructure eigenmodes from other partitions. Compared with typical CMS methods, our approach improves the accuracy by 3 orders of magnitude and achieves better strong and weak scaling in both time and memory cost. The advantages inherited from CMS, i.e. fast local updating and low communication cost between Message Passing Interface (MPI) ranks in large-scale distributed computing clusters are verified as well.
Eigenpair extractions are crucial for various applications in geometry processing and graphics. State of the Art libraries like ARPACK or Spectra rely on the implicitly restarted Lanczos iteration to extract eigenpairs efficiently. However for some large scale problems they lack convergence speed and robustness. In this paper we present a simple multigrid extension to accelerate the convergence and robustness of the implicitly restarted Lanczos method, and we demonstrate the efficiency of our method on a variety of problems commonly found in geometry processing and graphics.
We present HumanFlow, a unified flow-matching-based framework that enables high-fidelity and controllable full-body human image generation under diverse human-centric control conditions. Despite recent progress, controllable human image generation poses a fundamental challenge in balancing high visual fidelity with strict adherence to human-centric control conditions. HumanFlow formulates human image generation as a conditional flow-matching process with deterministic generation dynamics. To incorporate such human-centric control conditions into the pretrained model, we introduce a unified control framework with Control Encoder and Token-ControlNet. A Control Encoder maps diverse conditions into a unified latent representation that is spatially aligned with the image latent space. Token-ControlNet is a lightweight control network architecturally aligned with the FLUX double-stream design. To address accurate structural control over human bodies, we further propose the Human Topology Consistency Loss (HTCL). HTCL regularizes conditional flow matching by constraining generated human configurations to a union of statistically grounded topology manifolds defined by normalized bone ratios and joint angles. To support large-scale training and systematic evaluation, we construct MiCoGen, a multi-condition human image dataset comprising over one million full-body human images with aligned text descriptions and rich human-centric control conditions. Extensive quantitative and qualitative evaluations on the MiCoGen dataset show that HumanFlow consistently achieves improved structural consistency than the existing diffusion-based and flow-matching-based methods, while maintaining high visual fidelity.
Volumetric effects such as smoke, fire, dust, and explosions are central to Visual Effects (VFX) production and are commonly represented as sparse, high-resolution VDB/OpenVDB sequences with dynamic topology. Despite rapid progress in diffusion-based 3D generation, work on sparse volumetric sequences remains difficult to compare and reproduce, due to the lack of large-scale, well organized datasets and standardized evaluation protocols. In this paper, we introduce a 1-million-sample VFX sequence of VDB dataset with standardized preprocessing, consistent metadata, and protocol-ready splits, together with a reproducible benchmark suite for both static volume generation and sequence volumes generation. We further provide an end-to-end evaluation pipeline and a scalable diffusion training framework, enabled by our Atomic-Continuous prior, which addresses the distributional mismatch between vanilla diffusion models and the intrinsic sparsity of VDB data. Our release establishes a practical infrastructure for reproducible research and systematic progress tracking in sparse volumetric sequence generation. Project website: https://vfxdb-official.github.io/VfxDB/.
Realistic sound propagation is essential for immersion in a virtual scene, yet physically accurate wave-based simulations remain computationally prohibitive for real-time applications. Wave coding methods address this limitation by precomputing and compressing impulse responses of a given scene into a set of scalar acoustic parameters, which can reach unmanageable sizes in large environments with many source-receiver pairs. We introduce Reciprocal Latent Fields (RLF), a memory-efficient framework for encoding and predicting these acoustic parameters. The RLF framework employs a volumetric grid of trainable latent embeddings decoded with a symmetric function, ensuring acoustic reciprocity. We study a variety of decoders and show that leveraging Riemannian metric learning leads to a better reproduction of acoustic phenomena in complex scenes. Experimental validation demonstrates that RLF maintains replication quality while reducing the memory footprint by several orders of magnitude. Furthermore, a MUSHRA-like subjective listening test indicates that sound rendered via RLF is perceptually indistinguishable from ground-truth simulations.
Online 3D reconstruction from monocular image sequences is a challenging and ongoing research topic. 3D Gaussian Splatting (3DGS), leveraging its high-quality real-time rendering capability, empowers online 3D reconstruction to represent dense scenes with enhanced expressiveness, and thus holds great promise for a wide range of applications such as robotics and AR/VR. However, existing online 3DGS methods still suffer from some key challenges: fragile camera pose estimation due to the lack of global optimization, and low optimization efficiency in large-scale or long-sequence scenarios. To address these issues, we propose a robust and efficient online voxelized 3DGS reconstruction framework integrated with global Sim(3) optimization, which enables reliable camera tracking and efficient global loop closure for both camera poses and voxelized 3DGS. To accelerate the convergence of the voxelized 3DGS, we further introduce a color residual learning strategy, which not only boosts optimization speed but also enhances rendering quality. Extensive experiments on diverse indoor and outdoor datasets demonstrate that our method achieves state-of-the-art performance in both camera pose estimation accuracy and rendering quality, while retaining real-time efficiency. Additionally, we develop and deploy a real-world UAV-based active reconstruction system grounded on our proposed method, validating its robustness and generalizability for practical online 3D reconstruction tasks. Our code and data are available at https://github.com/TrickyGo/MoonSplat..
3D Gaussian Splatting (3DGS) has become the method of choice for reconstructing and real-time rendering of captured scenes. To capture a scene with good visual quality, continuous image sequences are usually combined with out-of-order shots for better scene coverage. Structure from motion can reconstruct such captures, but only after they are all available and often with high computational cost. Incremental reconstruction methods – often derived from SLAM solutions – provide immediate feedback, but cannot handle the out-of-order capture we require. We provide the first immediate feedback solution for such radiance field capture that provides global consistency. We first introduce a method for fast matching in out-of-order sequences, by repurposing visual place recognition models and a covisibility graph, and provide an efficient way to find highly connected keyframes, improving quality even for ordered sequences. We show how these steps – together with GPU optimization and careful Gaussian primitive placement – provide fast local reconstruction, in our challenging radiance field reconstruction case. We then introduce a novel cluster-based method, again using the covisibility graph, to provide efficient loop closure that does not require sequential input. Finally, to handle large scenes in our context, we introduce a progressive hierarchy that allows our method to scale to large environments, without compromising efficiency. Our results show we provide immediate feedback 3DGS reconstruction with good visual quality in several datasets, with up to thousands of input images.