Creative face stylization aims to render portraits in diverse visual idioms such as cartoons, sketches, and paintings while retaining recognizable identity. However, current identity encoders, which are typically trained and calibrated on natural photographs, exhibit severe brittleness under stylization. They often mistake changes in texture or color palette for identity drift or fail to detect geometric exaggerations. This reveals the lack of a style-agnostic framework to evaluate and supervise identity consistency across varying styles and strengths. To address this gap, we introduce StyleID, a human perception-aware dataset and evaluation framework for facial identity under stylization. StyleID comprises two datasets: (i) StyleBench-H, a benchmark that captures human same-different verification judgments across diffusion- and flow-matching-based stylization at multiple style strengths, and (ii) StyleBench-S, a supervision set derived from psychometric recognition-strength curves obtained through controlled two-alternative forced-choice (2AFC) experiments. Leveraging StyleBench-S, we fine-tune existing semantic encoders to align their similarity orderings with human perception across styles and strengths. Experiments demonstrate that our calibrated models yield significantly higher correlation with human judgments and enhanced robustness for out-of-domain, artist drawn portraits. All of our datasets, code, and pretrained models will be publicly available.
论文检索
输入标题、作者或关键词,从 326 篇学术成果中精准定位
End-to-end (e2e) latency in head-mounted displays (HMD) is the time delay between a physical change in the world (e.g., a user's head movement) and the moment the display updates to reflect that change. Tracking, rendering, and other computation in real systems invariably introduce some amount of e2e latency to all HMDs. In modern devices this latency is usually in the range of 12-60 milliseconds which is partially addressed through pose prediction and late stage reprojection which means that perceptual studies and user experience evaluations cannot explore latencies below these values. Here, we introduce a video passthrough HMD, called Camsicle, which is capable of 2-millisecond e2e latency and, additionally, uses a catadioptric design to achieve perspective-correct passthrough without reprojection. This platform enables naturalistic user studies to interrogate the impacts of latency on user experience, preference, and performance. Across two user studies and 57 participants we find that 2 and 14.3 millisecond latencies are preferred over 23 and 29 milliseconds when attempting to catch a ball. Additionally, we compare individual latency preferences in this naturalistic ball-catching task to psychophysical thresholds for latency detection in a reference-grade system with zero latency to investigate how psychophysical thresholds may relate to subjective evaluations in naturalistic scenarios.
JGS2 is a Jacobi-like GPU simulation algorithm. It avoids the overshooting issue by augmenting each subproblem with a perturbation subspace that predicts the global influence of the local solve. The efficiency of JGS2 is due to Cubature-based subspace integration at each subproblem. Being a data-driven method, Cubature requires a set of representative deformed poses that cover deformations likely to occur in the simulation. This requirement is unlikely for high-resolution deformation with rich local details. Therefore, simulation performance and convergence degenerate when Cubature extrapolates. This paper proposes a training-free subspace integration algorithm based on classic Gaussian quadrature (GQ). We leverage the fact that the subproblem's subspace bases can be well-approximated by a low-degree multivariable polynomial, which suggests GQ an excellent candidate for Cubature substitute. To this end, we introduce a novel algorithm that adaptively generates the integration region for each subproblem. As a result, GQ integration can be analytically retrieved without cumbersome data generation and training. We also show how to handle frictional contact by modifying the pre-computed perturbation subspace. The resulting JGS2-GQ framework is more versatile than the vanilla JGS2 method. It is more stable for large and novel deformations, and is free of data generation and expensive training, while maintaining a near second-order convergence that is comparable to Newton. Performance-wise, JGS2-GQ is as efficient as JGS2, which is three orders faster than classic CPU methods and up to two orders faster than classic GPU algorithms. When novel deformation occurs, JGS2-GQ outperforms JGS2 over 50%.
Lattice metamaterials enable lightweight, multifunctional structures, yet homogenization-based evaluation of their effective properties remains computationally expensive. Neural surrogates offer speed but often lack the accuracy and stability required for engineering-grade simulations. We introduce GMT, a Geometric Multigrid Transformer - a neural solver with high numerical fidelity for fast and reliable lattice homogenization. GMT achieves architectural alignment with Geometric Multigrid (GMG) by restructuring Point Transformer V3 to operate across sparse GMG hierarchies, capturing long-range dependencies and cross-level interactions essential for multigrid convergence. To enforce physical consistency, GMT incorporates physics-aware positional encoding for strict enforcement of periodicity and predicts both the finest-level solution and multi-level residual corrections. These predictions deliver a spectrally-aligned initialization, enabling end-to-end training under physics-informed and solver-aware losses and requiring only a single GMG V-cycle refinement to reach convergence. This fusion of neural prediction and numerical rigor achieves relative residual errors of 10-5 with a 160× speedup over state-of-the-art GPU-based solvers at equivalent accuracy - particularly at high resolutions (e.g., 5123), where traditional methods become most costly. We validate GMT across mechanical and thermal domains, demonstrate robust generalization to unseen geometries and non-periodic settings, and showcase scalability to high resolutions - enabling real-time design iteration, multi-scale simulations, high-throughput material discovery, and inverse design.
We introduce a high-performance framework for solving nonconvex optimization problems arising in simulation and geometry processing applications. Our framework especially benefits high-dimensional problems whose unknowns are spatial coordinates in ℝd and whose objectives are expressed as a sum of element energies over small stencils. The Hessians of these problems have a block-sparse structure, where the nonzero entries are grouped into dense d × d blocks. Many works have exploited this structure to speed up the Hessian-vector products and nodal smoothers of iterative solvers. In contrast, accelerating direct sparse linear solvers, preferred for applications requiring high accuracy on unstructured grids, has been less explored. Our framework leverages block structure to accelerate all phases of a Newton-type minimization algorithm, from sparsity pattern construction to system assembly, matrix factorization, and solves. A fundamental piece of the framework is BlockCatamari, our highly tuned, block-accelerated multifrontal sparse Cholesky code that achieves significantly faster factorizations than existing direct solvers on shared-memory systems. We combine this solver with our fast parallel Hessian assembly routine, heuristics for conditionally enabling per-element Hessian projections, and strategies for minimizing symbolic factorization re-computations. We show that this leads to dramatic speedups on a large suite of benchmark problems involving injective surface parametrization and elasticity simulation with contact, and we further demonstrate the ease with which new energies can be defined at the different levels of abstraction offered by our framework.
The theory of light transport on Gaussian process implicit surface (GPIS) provides a unified framework for rendering surfaces, participating media, and the intermediate spectrum. However, previous approaches rely on brute-force ray marching for surface intersections, requiring full noise evaluations at each marching point, whether using multivariate Gaussian sampling or sparse convolution noise approximation. This imposes a severe limitation on the rendering efficiency. In this paper, we derive bounds to significantly reduce the total number of full noise evaluations, leading to efficient ray marching for ray-surface intersections. We introduce stratified Bernoulli impulses, enabling a fast point-level bound for individual realizations to replace unnecessary full noise evaluations. To further reduce the number of point-level bound evaluations, we propose a region-level bound, leveraging a spatial acceleration structure to prune probabilistically empty regions, thereby avoiding unnecessary marching points in advance. By combining these two bounds, our bounded ray marching accelerates ray-surface intersections in GPIS, and consequently significantly improves overall GPIS rendering efficiency. Code for this paper are at https://github.com/Cchen-77/bounded-gpis.
We present a diffusion-based method for relighting dynamic portrait videos with photorealism and temporal consistency. Our method is fueled by a hybrid training dataset that consists of real-captured and rendered dynamic portrait videos with diverse subject appearances, facial motions, head poses, and known lighting conditions. Specifically, we construct an LED-based lighting system for realistic lighting emulation and high-speed video relighting data acquisition. By leveraging the image priors embedded in pre-trained video diffusion models, and using per-frame high dynamic range (HDR) environment map as lighting control, we train a high-performance generative model for realistic and identity-preserving dynamic portrait video relighting. In addition to the environment map control, our model uses a synthesized background image to enable control on the camera's exposure level and color tone. Our model can produce temporally consistent relit portrait video that looks realistic and harmonious under a provided new environment and faithfully preserve the subject's expression and fine facial features, including skin tone, wrinkles, and facial hair. Our model generalizes well to unseen data, in terms of the subject appearance, motion, and lighting condition. We perform extensive experiments on relighting in-the-wild videos with various environment maps and demonstrate practical applications on portrait photography. Results show that our method achieves state-of-the-art performance in photorealism, lighting harmony, and temporal consistency. Our project page: https://yufanzhang82.github.io/PixelCube/.
Being able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject-specific video diffusion-based framework for relighting full-body human performances in a temporally consistent way. Our model is trained on a hybrid dataset of pixel-aligned video relighting pairs, covering a diverse combination of lighting conditions, performances and viewpoints. To acquire such dataset, we combine traditional static One-Light-at-a-Time (OLAT) capture and a novel dynamic performance capture in which two smoothly varying lighting sequences are rapidly interleaved. Because the lighting operates above the human flicker-fusion threshold, the interleaving does not appears to strobe. We train our video relighting model from a pretrained text-to-video model to fully leverage the generative priors for producing high quality videos. To achieve accurate lighting control, we introduce a new lighting conditioning method that represents each light source as a token. We further condition on sequences of lighting using masked attention to support dynamic lighting control. Together with a carefully designed data augmentation pipeline, we achieve photorealistic, robust, and temporally consistent video relighting of subject-specific human performances.
We present Orbit-Space Geometric Probability Paths (OGPP), a particle-native flow-matching framework for generative modeling of particle systems. OGPP is motivated by two insights: (i) particles are defined up to permutation symmetries, so anonymous indexing inflates per-index target variance and yields curved, hard-to-learn flows; (ii) particles live in physical space, so the flow's terminal velocity has physical meaning and can encode geometric attributes (e.g., surface normals). OGPP instantiates three key components: (1) orbit-space canonicalization of the probability-path terminal endpoint, (2) particle index embeddings for role specialization, and (3) geometric probability paths with arc-length-aware terminal velocities that generate normals as a byproduct of the flow. We evaluate OGPP on minimal-surface benchmarks, where it reduces metric error by up to two orders of magnitude in a single inference step; on ShapeNet, where it matches the state-of-the-art with 5× fewer steps and reaches airplane EMD comparable to DiT-3D with 26× fewer parameters and 5× fewer steps; and on single-shape encoding, where it produces normals and reconstructions competitive with 6D generators while operating entirely in 3D.
We propose a unified, few-step generative modeling framework based on cumulative flow maps for long-range transport in probability space, inspired by flow-map techniques for physical transport and dynamics. At its core is a cumulative-flow abstraction that connects local, instantaneous updates with finite-time transport, enabling generative models to reason about global state transitions. This perspective yields a unified few-step framework built on cumulative transport and cumulative parameterization that applies broadly to existing diffusion- and flow-based models without being tied to a specific prediction instantiation. Our formulation supports few-step and even one-step generation while preserving synthesis quality, requiring only minimal changes to time embeddings and training objectives, and no increase in model capacity. We demonstrate its effectiveness across diverse tasks, including image generation, geometric distribution modeling, joint prediction, and SDF generation, with reduced inference cost.
Shape deformation for morphing or editing purposes is a central challenge in Computer Graphics. While numerous methods exist, few allow for the explicit evaluation of the deformation at arbitrary times and locations directly, without resorting to intricate advection or interpolation schemes. In this paper, we propose a method that provides an explicit expression of the deformation parameterized as a flow, for continuously deforming shapes defined implicitly. Implicit surfaces are indeed particularly well suited for deformation tasks, since they inherently account for both the surface and the enclosed volume. Our approach leverages invertible neural networks to ensure theoretically that the deformation is a valid flow, while also providing differential quantities useful for geometric regularization. We demonstrate applications of this flow to shape morphing with and without landmarks, shape editing, and pairwise-to-any morphing where we compute pairwise morphings to a canonical shape allowing to deduce transformations between any pair through flow composition. The code for our method is available at https://github.com/camillebnm/explicit_flows_for_implicit_surfaces.
We introduce a spatially discrete formulation of the Schrödinger bridge problem on meshes and grids that enables structure-preserving and scalable interpolation between probability distributions. Our approach builds on the duality between entropy-regularized optimal transport and the log-heat equation, deriving a discrete theory that is compatible with mesh-based finite element discretizations. The resulting Sinkhorn algorithm alternates application of the heat kernel with multiplicative updates to enforce marginal constraints. Compared to interpolation via Wasserstein barycenters, our formulation produces sharper interpolants for a given level of regularization and enforces exact endpoint marginals, in addition to enjoying faster computation. It also scales to high-resolution meshes and finer temporal discretizations, avoiding the prohibitive cost of directly discretizing dynamical transport. We demonstrate our approach across mesh- and grid-based applications, including displacement interpolation, shape interpolation, and color histogram manipulation, highlighting its ability to achieve geometric fidelity with computational efficiency.
College of Computing and Data Science, Nanyang Technological University, Singapore Multi-layer perceptrons (MLPs) are a standard tool for learning and function approximation, but they inherently produce globally smooth outputs. Consequently, they struggle to represent functions that are continuous yet intentionally non-differentiable (i.e., functions with prescribed C0 sharp features) without ad hoc post-processing. We present SharpNet, a modified MLP architecture that encodes user-specified sharp features by augmenting the network with an auxiliary feature function defined as the solution to Poisson's equation with jump Neumann boundary conditions. This feature function is evaluated via an efficient local integral and is fully differentiable with respect to the feature locations, allowing us to jointly optimize both the feature locations and the MLP parameters to recover the target function or geometry. This construction provides precise control over where non-differentiability occurs, enforcing the desired C0 behavior at feature locations while preserving smoothness elsewhere. We validate SharpNet on 2D problems and 3D CAD reconstruction, and compare it with several state-of-the-art baselines. In both settings, SharpNet accurately recovers sharp edges and corners while remaining smooth away from them, whereas existing methods tend to blur gradient discontinuities. Qualitative and quantitative results demonstrate the effectiveness of our approach. Our project page, code and models are publicly available at https://sharpnettech.github.io.
We propose a system for differentiating through solutions to geometry processing problems. Our system differentiates a broad class of geometric algorithms, exploiting existing fast problem-specific schemes common to geometry processing, including local-global and ADMM solvers. It is compatible with machine learning frameworks, opening doors to new classes of inverse geometry processing applications. We marry the scatter-gather approach to mesh processing with tensor-based workflows and rely on the adjoint method applied to user-specified imperative code to generate an efficient backward pass behind the scenes. We demonstrate our approach by differentiating through mean curvature flow, spectral conformal parameterization, geodesic distance computation, and as-rigid-as-possible deformation, examining usability and performance on these applications. Our system allows practitioners to differentiate through existing geometry processing algorithms without needing to reformulate them, resulting in low implementation effort, fast runtimes, and lower memory requirements than differentiable optimization tools not tailored to geometry processing.
Capturing cloth motion with high fidelity is challenging due to fine-scale wrinkles, large deformation, and frequent self-occlusion. We present a 4D (spatio-temporal) cloth capture system that achieves 1 mm spatial resolution using only 16 RGB cameras. Our approach uses a two-level marker pattern: sparse, colored L-shaped markers provide robust detection and orientation, while dense noise patterns within each marker enable both marker identification and precise keypoint localization. By unwarping detected markers to a canonical frame, we factor out perspective distortion and most of cloth deformation, allowing the localizer to achieve sub-pixel accuracy. The localized keypoints are triangulated across views to form an incomplete point cloud. A physics-based optimization then deforms a template mesh to match the captured geometry while maintaining penetration-free constraints and physical plausibility for occluded regions. Our method produces temporally coherent sequences that faithfully capture fine wrinkles and folds even during complex motions with self-contact.
Volumetric video enables immersive free viewpoint rendering of dynamic real world scenes, yet existing methods struggle with long sequences and complex motions, often leading to temporal instability and visual artifacts. To address these challenges, we propose ATGS, a Gaussian splatting based framework for volumetric video reconstruction. Our key insight is that explicitly tracking long term complex motion with individual Gaussian primitives is inherently unstable. Instead, we organize Gaussians around time conditioned anchors that localize their spatial and temporal support, thereby reducing long range motion complexity. We further introduce a temporal windowing strategy to activate only anchors relevant to the queried time, which improves scalability and temporal coherence. In addition, to ensure spatial and temporal stability, we design a compact set of multi level anchor features that encode global features, local spatial features, and local temporal features, jointly constraining Gaussian generation. Extensive experiments demonstrate that ATGS consistently outperforms prior methods on long sequence volumetric videos with complex motions. Project page: https://github.com/WuJH2001/ATGS.
The rapid adoption of volumetric capture technologies has created a pressing need for efficient storage and streaming of dynamic 3D content. Unfortunately, current compression standards often treat dynamic sequences as independent frames or rely on computationally expensive non-rigid registration, making them unsuitable for real-time applications or large-scale environments. In this paper, we present a novel end-to-end compression framework for dynamic Truncated Signed Distance Field volumes derived from captured 3D content, leveraging a representation that is temporally stable, easily parallelizable, and already widely used in scene reconstruction and volumetric fusion pipelines. We then adapt classic 2D video coding paradigms such as spatial coding via Discrete Cosine Transform and temporal coding using a real-time motion compensation pipeline to provide robust, real-time, and training-free encoding and decoding for 3D content. Extensive evaluations on human performance captures demonstrate that our codec achieves ~35% bitrate savings at equal distortion while operating in real time at 30 FPS, while stronger temporal coherence in large-scale synthetic environments yields up to 12× bitrate reduction at equal distortion.
Articulation modeling aims to infer movable parts and their motion parameters for a 3D object, enabling interactive animation, simulation, and shape editing. In this paper, we present Sketch2Arti, the first sketch-based articulation modeling system for CAD objects. Our key observation is that designers naturally communicate articulation intent through lightweight sketches (e.g., arrows and strokes) that indicate how parts should move, yet translating such sketches into articulated 3D models remains largely manual. Sketch2Arti bridges this gap by enabling users to specify articulation through simple 2D sketches drawn from a chosen viewpoint. Given a CAD model and user sketches, our approach automatically discovers the corresponding movable parts and predicts their motion parameters, allowing iterative modeling of multiple articulations on complex objects with fine-grained control. Importantly, Sketch2Arti is trained in a category-agnostic manner without requiring object category information, leading to strong generalization to diverse objects beyond existing articulation datasets. Moreover, for shell models lacking interior structures, Sketch2Arti supports controllable internal completion guided by user sketches, generating plausible internal components consistent with the existing geometry and predicted motion constraints. Comprehensive experiments and user evaluations demonstrate the effectiveness, controllability, and generalization of Sketch2Arti. The code, dataset, and the prototype system are at https://arlo-yang.github.io/Sketch2Arti.
Stochastic Monte Carlo solvers for partial differential equations (PDEs) recently gained popularity in computer graphics, finding applications in geometry processing, rendering, simulation, and visualization. At present, there exists no Monte Carlo solver for the rendering of biharmonic diffusion curves, an artist-friendly smooth vector graphics primitive. The fourth-order biharmonic equation of biharmonic diffusion curves can be split into two second-order PDEs, namely a Laplace and a Poisson equation. However, since biharmonic diffusion curves set Dirichlet and inhomogeneous Neumann conditions at the same time, these two second-order PDEs are tightly coupled and can hence not be solved directly. We propose to treat the rendering of biharmonic diffusion curves as an inverse problem, in which the Dirichlet data of the Laplace equation is unknown. We formulate a variational energy optimization, such that the user-defined boundary conditions are met. Thereby, the necessary gradients are estimated stochastically by solving two second-order problems with Dirichlet boundary conditions only.
This paper describes a strategy for vectorizing 3D scenes with proper occlusion. Given a collection of curves derived from 3D geometry (silhouettes, isolines, material boundaries, etc.), we produce a 2D vector image that partitions the image plane into a planar map of solid shaded regions. The method is agnostic to surface representation, handling for instance curves obtained from polygonal, NURBS or subdivision surfaces. The output likewise supports curve segments of arbitrary parametric type. Our key observation is that the spatial hierarchy used to accelerate curve-curve intersections provides the fundamental representation of the planar map itself. This approach provides geometric flexibility, since general curve-curve intersection problems are replaced with simpler curve-line intersection. Simultaneously, it provides robustness since cells of the spatial hierarchy define a well-defined planar map, even in the presence of numerical errors. For instance, it automatically handles "curve soup" where segment endpoints are not explicitly connected in the input file. The method scales to a large number of primitives, and is output sensitive: it resolves intersections only up to a user-defined precision—while still providing topologically valid output. We evaluate the method on a collection of challenging tests, showing that it is both more robust and orders of magnitude more efficient than existing curve arrangement techniques, such as those found in CGAL.