论文检索

输入标题、作者或关键词,从 100,903 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
100,903篇论文
第 12 / 5046 页

Ling-Hao Chen, Zixin Yin, Duomin Wang, Xianfang Zeng, Gang Yu

This work explores the motion transfer from one video to another, which is crucial in animation for diverse characters. Previously, video motion transfer has been largely explored between human and human-like characters, enabling a lot of applications in digital creation. However, these approaches encounter a main limitation. Specifically, related technical pipelines heavily rely on a predefined human skeleton structure and accordingly require skeleton-conditional model training. On the one hand, these methods are difficult to generalize to diverse characters, such as animals from different species, while preserving their unique motion styles. On the other hand, labeled data in diverse skeletons is limited, which additionally restricts the large-scale training for the task. In this paper, we jump out of the skeleton-based motion transfer framework and propose a training-free motion transfer framework, named Motion4Motion. Motion4Motion models the motion flow of the character in a video instead of skeletons, which makes motion transfer across species easier. Extensive experimental results and novel applications show our methods outperform baselines impressively.

Nadav Z. Cohen, Ofir Abramovich, Ariel Shamir

Text-to-image diffusion models generate images by gradually converting white Gaussian noise into a natural image. White Gaussian noise is well suited for producing diverse outputs from a single text prompt due to its absence of structure. However, this very property limits control over, and predictability of, specific visual attributes, as the noise is not human-interpretable. In this work, we investigate the characteristics of the input noise in diffusion models. We show that, although all frequencies in white Gaussian noise have comparable statistical energy, low-frequency components primarily determine the image’s global structure and color composition, while high-frequency components control finer details. Building on this observation, we demonstrate that simple manipulations of the low-frequency noise using low-frequency image priors can effectively condition the generation process to reconstruct these low-frequency visual cues. This allows us to define a simple, training-free method with minimal overhead that steers overall image structure and color, while letting high-frequency components freely emerge as fine details, enabling variability across generated outputs.

Alvin Shi, Florence Bertails-Descoubes, A. M. Darke, Theodore Kim

Due to its highly curved geometry, tightly coiled hair is challenging to model and edit using standard position-based tools. In this work we propose using material curvatures and twists to analyze and edit tightly coiled hair styles. Our method relies on the geometry of super-helices, primitives parametrized by piecewise constant curvatures and twists, whose helical geometry naturally resembles a coiled hair strand. Using this curvature/twist space, we introduce new editing tools that allow us to expand, shrink, “ruffle”, interpolate or guide the position of coiled hair in a natural way. We present analytical expressions for geometry and gradients that allow our method to run efficiently and without the need for any training data. We successfully apply our tools to highly coiled simulated hairs, as well as those generated procedurally.

Yuefan Shen, Yican Dong, Xiufeng Huang, Zhongtian Zheng, Youyi Zheng, Kui Wu 0003

The fundamental limitation of traditional strand-based modeling is not simply data scarcity, but the ill-posedness of inferring complex 3D fields from 2D imagery without structural constraints. This unconstrained regression leads to catastrophic failures in resolving both global occlusion (e.g., in ponytails) and local directionality (e.g., in curls), resulting in over-smoothed, plausible-but-incorrect geometries. To resolve this, we integrate the strong geometric priors of Large Reconstruction Models (LRMs) into the strand generation pipeline. Using the LRM mesh as a structural anchor, we employ a novel Dual Orientation AutoEncoder to lift coarse geometry into high-fidelity strands. By resolving vector field singularities through latent-space optimization and surface-guided refinement, our method effectively disentangles complex topological structures, setting a new benchmark for robustness and accuracy in hair reconstruction.

Pengpei Hong, Song Zhang 0007, Daqi Lin, Markus Kettunen 0001, Chris Wyman, Cem Yuksel

ReSTIR [Bitterli et al. 2020] efficiently reduces path tracing noise by reusing samples spatiotemporally. Recent ReSTIR methods [Liu et al. 2025] improve temporal reuse from prior frame primary hits. But disocclusions still invalidate some pixel histories, degrading quality with spatially varying noise. We address disocclusions in ReSTIR by introducing multiple screen-space layers, using reservoir splatting to shift samples between layers. This reuses previously occluded samples that become visible again, reducing noise in disoccluded regions. While shifting samples across layers typically requires tracing multiple rays, we introduce depth ranges and redefine the integration domain to eliminate many ray queries, improving performance. To minimize costs for maintaining many layers, we selectively track only active domains propagated from previous frames. We validate across multiple scenes, showing greatly reduced disocclusion noise while incurring only a small incremental cost.

Yu-Chen Wang, Markus Kettunen 0001, Daqi Lin, Chris Wyman, Lifan Wu, Shuang Zhao

Rendering complex scenes in real time remains challenging due to strict performance constraints. Geometric level of detail (LoD) is widely used to reduce cost by replacing high-frequency geometry with prefiltered representations. ReSTIR significantly improves real-time rendering quality by reusing samples across space and time; however, its effectiveness degrades in the presence of complex geometry, where high-frequency detail reduces reuse efficiency. Moreover, prior ReSTIR methods require the same mesh topology across frames, causing sample reuse to break when switching LoD. In this paper, we introduce a surface point mapping that enables sample reuse across frames containing meshes with different topologies. Building on our method, we enable LoD-aware ReSTIR by maintaining valid spatiotemporal reuse under geometry LoD changes. Our approach restores reuse efficiency in LoD scenes and significantly improves rendering quality compared to prior ReSTIR methods.

Zhong Shi, Cunhao Wu, Lifan Wu, Kun Xu 0003

Real-time path tracing demands high visual quality under extremely tight sampling budgets, often relying on reservoir-based spatio-temporal importance resampling (ReSTIR) to maximize sample quality. However, ReSTIR typically estimates the pixel integral using a single representative sample selected via a scalar target function (e.g., luminance). This inevitably leads to color noise in scenes with complex chromatic lighting or materials. In this work, we present Reservoir-based Spatio-Temporal Control Variates (ReSTCV), a novel framework that addresses this problem by integrating Spatio-Temporal Control Variates (STCV) into ReSTIR. We revisit image-space control variates—originally an offline technique—and adapt them for real-time rendering by spatio-temporal sample reuse. This unified approach combines the benefits of both techniques, enabling us to suppress color noise while maintaining the efficiency of ReSTIR. Our method introduces minimal computational overhead and requires only minor modifications to existing ReSTIR pipelines. We demonstrate that ReSTCV produces significantly cleaner images with stable colors across a variety of dynamic scenes, marking the first practical application of spatio-temporal control variates in real-time path tracing.

Cem Yuksel

Cubic polynomial curve modeling with Bézier handles is ubiquitous, but it provides no practical mechanism for achieving curvature continuity. We present two distinct solutions for this problem. The first one is continuity-enhancing degree elevation that offers a general solution to achieve any level of continuity by converting the given curve to a polynomial of a higher degree. Our second solution is continuity-enhancing splits, which is specific to cubic curves and achieves curvature continuity by splitting the curve pieces but maintaining the piecewise cubic polynomial form. Both of these solutions utilize a local optimization process with a closed-form solution, achieving continuity enhancement with a constant computation overhead per piece. We also explain how to incorporate linearity constraints to seamlessly form linear curve pieces, when desired. Our solutions are effective in extending the popular curve modeling interface with Bézier handles to splines with curvature (or higher) continuity. Furthermore, we show that our solutions can also be used for defining new interpolating curve formulations with desirable properties, and they can be used with higher-dimensional curves or surfaces.

Bowen Zheng, Linjun Wu, Xinwei Jiang, Yujin Chai, Zijiao Zeng, He Wang 0002, Xiaogang Jin 0001

Motion warping is a core technique in character animation that enables the adaptation of existing motion data to novel spatio-temporal constraints. Conventional motion warping methods often rely on heuristic modifications that can violate physical consistency or introduce visual artifacts. More recent learning-based editing approaches improve realism, but many of them encode motion into tightly entangled latent space, which makes them struggle to balance editing flexibility and content preservation. To address this, we propose a novel deep motion warping framework that explicitly disentangles the motion structure from global and stylistic attributes for intuitive motion editing. Our key insight is to leverage learned phase features as a continuous and robust representation of the underlying structure, and explicitly disentangle motion into root velocity, phase, and learned latent variables using a phase-conditioned diffusion autoencoder. This design supports a wide range of editing operations, including root motion warping, motion exaggeration, time warping, and style transfer by directly manipulating decoupled components, without requiring paired training data. Extensive experiments demonstrate that our approach enables high-level, flexible motion editing while strictly preserving the structural consistency and physical plausibility of the source motion

Yi Shi 0008, Yifeng Jiang 0009, Chen Tessler, Xue Bin Peng

Developing controllers capable of completing a wide range of tasks in a natural and life-like manner is a key challenge in enabling practical applications of physics-based character animation. In this work, we introduce Generative Pretrained Controllers (GPC), which leverage tokenization and next-token modeling to create general-purpose, reusable generative controllers from large-scale motion datasets. Our framework utilizes end-to-end reinforcement learning to jointly optimize a "motion vocabulary", modeled via Finite Scalar Quantization (FSQ), along with a corresponding control policy that can map the discrete codes to physics-based controls. After the "codebook" has been learned, the underlying structure of this large vocabulary is modeled by training a GPT-style autoregressive transformer, leading to a powerful generative controller that generates controls for a physically simulated character by performing next-token prediction. Once the generative controller has been trained, we propose a suite of adaptation techniques for finetuning the controller for new downstream tasks. Our proposed framework greatly simplifies the training process compared to previous tokenized methods, and achieves a 99.98% success rate in reproducing a vast corpus of motion clips. The generative controller exhibits a variety of natural emergent behaviors, such as responsive behaviors to perturbations and recovery behaviors after falling. This results in highly robust general purpose controllers for a variety of downstream applications.

Xianyao Zhang, Gerhard Röthlin, Tunç Ozan Aydin, Farnood Salehi, Marios Papas

Deep images store a variable number of “bins” within each pixel, enabling deep compositing workflows with clean separation of overlapping objects, but Monte Carlo noise limits their practical use. Existing deep image denoisers are limited in quality, generality, and temporal processing because ragged bin neighborhoods are incompatible with fixed convolution kernels. We introduce the ragged neighborhood attention (RaNA) operator, which extends neighborhood attention to semi-structured deep images by dynamically resolving bin-to-bin relationships across spatial and temporal neighborhoods. Using RaNA, we build the first spatiotemporal neural denoiser for deep Monte Carlo renderings, termed RaNAD, with a multi-U-Net backbone, multi-scale reconstruction, and an optimized CUDA implementation. RaNAD handles both deep-Z and deep-OID variants and supports temporal windows of up to 7 frames. Compared with the previous state of the art for deep image denoising, RaNAD improves denoising quality substantially while preserving the layered structure needed for deep compositing; when its output is flattened for evaluation, it also attains quality competitive with strong flat image denoisers. As a kernel-based method, RaNAD efficiently denoises multi-AOV images, and temporal processing further improves quality and stability, making the method practical for offline production rendering and compositing workflows.

Bing Xu, Mukund Varma T., Cheng Wang, Tzu-Mao Li, Lifan Wu, Bartlomiej Wronski, Ravi Ramamoorthi, Marco Salvi

Global illumination (GI) is essential for realism but remains computationally expensive. While per-scene neural methods lack generalization and screen-space approaches inherently suffer from view inconsistency, prior 3D neural rendering methods face a severe scalability barrier, restricting them to small, object-centric meshes. To overcome these trade-offs, we introduce a generalizable light transport 3D embedding that predicts global illumination directly from 3D scene configurations without rasterized or path-traced illumination cues, per-scene retraining or screen-space limitations. We employ a point-based representation to decouple our embedding from the original scene topology, then utilize a linear-complexity transformer to encode long-range light transport. This design scales to environments with millions of triangles, enabling the first generalizable GI learning on complex, high-fidelity indoor scenes, far beyond prior limits. To achieve this, we enforce a local query mechanism where rendering queries are processed independently under 3D supervision. This ensures constant complexity per pixel relative to scene size, yielding view-consistent and resolution-agnostic rendering without the memory bottlenecks typical of globally coupled attention. We further demonstrate versatility by re-targeting the encoder with limited fine-tuning, presenting preliminary results on spatial-directional radiance field prediction for glossy materials and validating transfer to downstream rendering tasks.

Howard Xiao, Jan Ackermann, Boyang Deng, Gordon Wetzstein

Ultra-high-resolution image sensors offer the potential to capture fine spatial details critical for many visual perception tasks, but acquiring and processing all pixels at full resolution is often infeasible under realistic bandwidth, latency, and power constraints. Existing approaches address this challenge through acquisition strategies such as spatial or temporal downsampling, which irrevocably discard information before task relevance can be assessed. In this work, we introduce a real-time, predictive, and task-aware foveated imaging system that operates directly at image acquisition time. Leveraging emerging dual-stream sensor architectures, our method dynamically allocates limited pixel bandwidth to task-relevant regions of interest while maintaining a low-resolution global context. We formulate foveated acquisition as a sensor attention policy–learning problem, in which past observations guide actions that determine future measurements, closing the perception–acquisition loop. Through extensive simulation across multiple perception tasks, we demonstrate that our approach achieves high task performance under strict pixel budgets and significantly outperforms relevant baselines operating at the same bandwidth. We further validate our system on a 200-megapixel dual-stream sensor, capturing real-world videos under realistic bandwidth and latency constraints, demonstrating the practical feasibility of task-driven, acquisition-time foveated imaging. Our project website is at https://howardxiao.ca/foveated/.