论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 41 / 81 页

Huabin Wang, Yingfan Cheng, Wu Zheng, Jiayuan Cheng, Xin Li 0248, Min Li 0033, Fei Liu

Near-infrared transmission through the finger can capture the vein structure for identity recognition. However, in outdoor applications, finger vein imaging is significantly affected by environmental illumination resulting in low recognition performance. Existing methods typically address this issue by constructing multi-illumination models, but collecting multi-illumination images from individual is challenging, and overexposure can cause venous structure distortion. This paper proposes MDA-Net, a Multi-illumination Domain Adaptive Network for finger vein recognition, which is engineered to excel in the dynamic outdoor lighting landscape with various conditions including overexposure, using only data collected under a single illumination for training. Firstly, an Illumination Feature Separation Network(IFSNet) is used to remove the illumination components and obtain illumination-invariant features; Then an Absorption Difference Feature Extraction network(ADFENet) is used to reduce the impact of venous structure distortion under illumination conditions, especially overexposure. To replicate the entire range from low-light to overexposure in outdoor scenarios, a novel Multi-Illumination Finger Vein Dataset (MIFVD) is constructed with significant illumination variations. Experimental results show that MDA-Net significantly improves recognition performance under complex illumination conditions, achieving a state-of-the-art (SOTA) average recognition rate of 91.67% and an average equal error rate (EER) of 0.96%. Further validation on public datasets SDU and USM, demonstrates SOTA EERs of 0.16% and 0.10%, respectively. The License for MIFVD can be accessed at: https://github.com/AHU-MedImagingIJR/MIFVD.

Ziwei Niu, Shiao Xie, Ziyue Wang 0005, Yen-Wei Chen 0001, Yueming Jin, Lanfen Lin

Single-source domain generalization (SDG) in medical image segmentation is a challenging yet practical task that efficiently enhances generalization ability while avoiding high annotation costs and privacy concerns. In this paper, we propose EIR-SDG, a novel SDG approach that explores domain-invariant representation for medical image segmentation. The core of EIR-SDG lies in mitigating the effect of style in the encoder while facilitating robust segmentation in the decoder. Concretely, we design a training-free texture and style diversity module that transforms images into diverse random appearances without requiring optimization or gradient updates, which simulates unseen target distributions while mitigating overfitting to regular patterns in synthetic data. Building on this, we devise a feature adaptive whitening module, which disentangles and whitens the style-sensitive feature correlations between original and augmented pairs, encouraging the encoder to learn invariant representations. Moreover, to facilitate robust segmentation in the decoder, a semantic representation optimization strategy is devised to enhance invariant representations by constraining the correlation between class prototypes to be consistent while improving segmentation boundary distinction by separating different class prototypes. Experiments on cross-modality abdominal, cross-sequence cardiac and cross-center prostate segmentation tasks demonstrate that our method achieves promising generalization capacity and outperforms the SOTA methods.

Zikai Zhang 0004, Xu Zhang 0085, Ziyi Li, Yidong Li, Yuanzhouhan Cao

Multimodal learning integrates diverse modalities to enhance robustness, yet real-world scenarios suffer from heterogeneous imbalance phenomena (noise interference, modality partial missing, intermodal information disparities), degrading performance through biased feature representations. Existing methods fail to adaptively modulate models under dynamic imbalance conditions. We propose GMML, a framework dynamically balancing multimodal gradients to counteract imbalance-induced biases: i) An imbalance-aware gradient modulation adaptively identifies contributions with smooth weight transitions to balance conflicting gradients; ii) A parameter constraint method enforces ℓ2-norm constraints on encoders, suppressing parameter oscillations and blocking noisy updates under modality missing/noise. Theoretically, GMML achieves a larger certified radius upper bound for complex imbalances, with convergence radius analysis providing theoretical guarantees. Experiments demonstrate superior robustness against three imbalance types, outperforming state-of-the-art by 3.3% and 2.3% in accuracy on KS and UCF-101 benchmarks. series Code: https://github.com/zhangzikai-security-ML/GMML.

Xuanchen Wang, Heng Wang 0007, Weidong Cai 0001

Modern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches often fail to produce high-quality dance videos that harmonize with both musical rhythm and user-defined choreography styles, limiting their applicability in real-world creative contexts. To address this gap, we introduce ChoreoMuse, a diffusion-based framework that uses SMPL format parameters and their variation version as intermediaries between music and video generation, thereby overcoming the usual constraints imposed by video resolution. Critically, ChoreoMuse supports style-controllable, high-fidelity dance video generation across diverse musical genres and individual dancer characteristics, including the flexibility to handle any reference individual at any resolution. Our method employs a novel music encoder MotionTune to capture motion cues from audio, ensuring that the generated choreography closely follows the beat and expressive qualities of the input music. To quantitatively evaluate how well the generated dances match both musical and choreographic styles, we introduce two new metrics that measure alignment with the intended stylistic cues. Extensive experiments confirm that ChoreoMuse achieves state-of-the-art performance across multiple dimensions, including video quality, beat alignment, dance diversity, and style adherence, demonstrating its potential as a robust solution for a wide range of creative applications. Video results can be found on our project page: https://choreomuse.github.io.

Yu Tong, Weihai Lu, Xiaoxi Cui, Yifan Mao, Zhejun Zhao

Lately, the academic community has been showing growing interest in multi-domain fake news detection, and in particular, incorporating multimodal information into this field has emerged as a highly promising research direction. However, existing methods often struggle with: (1) Insufficient intrinsic domain adaptation during representation Learning; (2) Amplified negative transfer from entangled domain style and content representations; and (3) Neglecting domain-varying modality uncertainty. To address these issues, we propose Domain-Aware Prompt Tuning (DAPT), an innovative framework for multimodal multi-domain fake news detection. DAPT leverages Multimodal Prompt Tuning for parameter-efficient domain adaptation of pretrain models. An Adaptive Domain Debias Module disentangles domain features from veracity signals guided by content to mitigate negative transfer. Furthermore, inspired by the Variational Information Bottleneck, an Uncertainty-Aware Multimodal Fusion mechanism adaptively aggregates modalities based on domain-specific reliability. Extensive experiments demonstrate that DAPT significantly outperforms state-of-the-art baselines on benchmark datasets.

Tianshun Han, Benjia Zhou, Ajian Liu 0001, Yanyan Liang 0001, Du Zhang, Zhen Lei 0001, Jun Wan 0001

Speech-driven 3D facial animation aims to synthesize realistic emotional facial expressions that match the input speech. However, existing approaches are constrained by two key limitations: (1) These methods rely on pre-trained models (e.g., Wav2Vec 2.0) as audio emotion feature extractors, which neglect critical frequency-domain characteristics, thereby emphasizing the challenge of discriminating between similar emotion categories. (2) They treat audio emotions as generic categorical states, ignoring individual differences in emotional expression, ultimately producing over-smoothed emotional representations that appear repetitive and stereotypical. To that end, we introduce PESTalk, a novel approach that generates 3D facial animations with Personalized Emotional Styles directly from speech inputs, thus significantly enhancing the realism of facial animations. Specifically, since acoustic frequency cues contain essential emotional information, we first propose a Dual-Stream Emotion Extractor (DSEE ), which captures both time-domain variations and frequency-domain characteristics of audio signals to extract fine-grained affective features and subtle emotional nuances. Furthermore, we design an Emotional Style Modeling Module (ESMM ) to achieve personalized emotional styles. This module first establishes a baseline representation for each subject based on voiceprint characteristics, then progressively refines it by continuously integrating emotional features. Ultimately, this process constructs a personalized emotional style representation for each subject in each emotion category, capturing their unique expression patterns. Finally, considering the scarcity of the 3D emotional talking face data, we employ an advanced facial capture model to extract pseudo facial blendshape coefficients from 2D emotional data, thereby constructing a large-scale 3D emotional talking face dataset with diverse emotions and personalized expressions (3D-EmoStyle). Extensive quantitative and qualitative evaluations show that PESTalk can generate realistic 3D facial animation and outperform state-of-the-art methods. The codes and dataset are available at: https://github.com/tianshunhan/PESTalk.

Kerun Mi, Guoliang Kang, Guangyu Li, Lin Zhao 0003, Tao Zhou 0002, Chen Gong 0002

Class-Incremental Unsupervised Domain Adaptation (CI-UDA) aims to adapt a model from a labeled source domain to an unlabeled target domain, where the sets of potential target classes appearing at different time steps are disjoint and are subsets of the source classes. The key to solving this problem lies in avoiding catastrophic forgetting of knowledge about previous target classes during continuously mitigating the domain shift. Most previous works cumbersomely combine two technical components. On one hand, they need to store and utilize rehearsal target sample from previous time steps to avoid catastrophic forgetting; on the other hand, they perform alignment only between classes shared across domains at each time step. Consequently, the memory will continuously increase and the asymmetric alignment may inevitably result in knowledge forgetting. In this paper, we propose to mine and preserve domain-invariant and class-agnostic knowledge to facilitate the CI-UDA task. Specifically, via using CLIP, we extract the class-agnostic properties which we name as ''attribute''. In our framework, we learn a ''key-value'' pair to represent an attribute, where the key corresponds to the visual prototype and the value is the textual prompt. We maintain two attribute dictionaries, each corresponding to a different domain. Then we perform attribute alignment across domains to mitigate the domain shift, via encouraging visual attention consistency and prediction consistency. Through attribute modeling and cross-domain alignment, we effectively reduce catastrophic knowledge forgetting while mitigating the domain shift, in a rehearsal-free way. Experiments on three CI-UDA benchmarks demonstrate that our method outperforms previous state-of-the-art methods and effectively alleviates catastrophic forgetting. Code is available at https://github.com/RyunMi/VisTA.

Chengcheng Xing, Yanyu Xu 0001, Yonghui Xu, Lizhen Cui 0001

Unified Anomaly Detection (UAD) aims to identify anomalies across diverse domains without access to target domain data during training. Unlike traditional anomaly detection methods that rely on training separate models for each domain, UAD employs a single model to generalize across multiple categories. A key challenge lies in the domain shift between seen and unseen data, which requires capturing invariant discriminative patterns between reference and query images across different domains during in-context learning for unified anomaly detection. To tackle this, we propose a novel UAD framework to learn the invariant discriminative patterns through pre-, in- and post-processing modules. First, a pre-processing VLM-guided data augmentation module generates diverse and semantically consist images, followed by a latent-space filtering mechanism. Second, an in-processing Adaptive VQ memory module stores representative discriminative patterns to enable robust residual comparison. Third, a post-processing GUR (Geometric distributions Upgrade Representation) feature augmentation module models geometric feature distributions to synthesize informative prompts, improving the quality of feature delta estimation for anomaly scoring. Extensive experiments on benchmark datasets demonstrate that our method achieves superior generalization in detecting anomalies across unseen domains, outperforming existing state-of-the-art approaches.

Juan Zhao 0007, Yudao Sun, Zhihai Yang, Cai Xu, Hongji Chen 0005, Fan Zhang 0112, Jianxin Li 0001

Deep neural networks on cloud platforms face growing security threats, with AI services increasingly relying on heterogeneous models for the same task to meet diverse user needs. Existing methods fail to distinguish benign modifications from malicious attacks in cross-model scenarios. To address this challenge, we propose a non-intrusive cross-model watermarking method that generates discriminative samples as universal keys, enabling authentication without altering model parameters or architectures. Specifically, we introduce a margin enhancement loss to amplify confidence gaps between benign and malicious behaviors, ensuring high transferability across models. Both theoretical analysis and experimental results demonstrate the high efficacy of our proposed method. The generated samples maintain high visual fidelity (SSIM > 0.99), achieve over 3 times higher discriminability than existing methods, retain over 93% accuracy under benign modifications, and detect malicious attacks with accuracy dropping below 9%. Overall, our proposed method provides a robust, transferable, and non-intrusive solution for cross-model authentication, making it ideal for real-world applications where security is critical.

Ruoxuan Zhang, Bin Wen 0001, Hongxia Xie, Yi Yao, Songhan Zuo, Jian-Yu Jiang-Lin, Hong-Han Shuai, Wen-Huang Cheng

Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusion models have shown strong capabilities in text-to-image generation, they struggle to handle structured multi-step scenarios like recipe illustration. Additionally, current recipe illustration methods are unable to adjust to the natural variability in recipe length, generating a fixed number of images regardless of the actual instructions structure. To address these limitations, we present CookAnything, a flexible and consistent diffusion-based framework that generates coherent, semantically distinct image sequences from textual cooking instructions of arbitrary length. The framework introduces three key components: (1) Step-wise Regional Control (SRC), which aligns textual steps with corresponding image regions within a single denoising process; (2) Flexible RoPE, a step-aware positional encoding mechanism that enhances both temporal coherence and spatial diversity; and (3) Cross-Step Consistency Control (CSCC), which maintains fine-grained ingredient consistency across steps. Experimental results on recipe illustration benchmarks show that CookAnything performs better than existing methods in training-based and training-free settings. The proposed framework supports scalable, high-quality visual synthesis of complex multi-step instructions and holds significant potential for broad applications in instructional media, and procedural content creation. More details are at https://github.com/zhangdaxia22/CookAnything.

Ru Jia, Xiaoqian Liang, Xubin Duan, Jianji Wang 0001, Nanning Zheng 0001

Despite recent advances in dynamic scene reconstruction, challenges from imbalanced camera distribution and inaccurate pose estimation in real-world datasets still persist, undermining the spatiotemporal consistency of reconstruction. In this paper, we propose HybridPlane, a novel representation that leverages the complementary advantages of cylindrical and Cartesian coordinate systems to achieve high-quality dynamic scene synthesis. Unlike Cartesian projection, which shares identical features in symmetric regions, cylindrical projection explicitly disentangles features from different viewpoints, thereby improving robustness against imbalanced camera distributions. Moreover, the synergy between these two coordinate systems in both projection and representational capacity enhances the model's ability to capture complex motions and fine-grained details. We further adopt the dynamic positional encoding strategy to enhance the smoothness of temporal interpolation under inaccurate camera poses by progressively regulating high-frequency signals without incurring additional computational overhead. Extensive experiments demonstrate that our versatile representation can be seamlessly integrated into various rendering pipelines, outperforming the previous methods in reconstruction quality while reducing computational and memory costs by approximately one-third.

Xiang Huang 0004, Ao Luo, Xiao Wu 0001, Zhaoquan Yuan

Human-Object Interaction (HOI) detection serves a broad spectrum of applications. Despite significant progress, current approaches encounter difficulties in effectively handling Non-Contact Human-Object Interaction (NCHOI) scenarios, where humans and objects remain physically apart. To address these challenges, this paper proposes a novel approach, named Latent Interactiveness Field Modeling (LIFM), which enhances HOI detection by capturing long-range contextual dependencies. Specifically, the Latent Interactiveness Field (LIF) is introduced to define potential interactive relationships between humans and objects. To complement this, the LIF Fusion Encoder is designed to adaptively fuse visual features with LIF, resulting in more informative and discriminative feature representations. The Mobile Scanning HOI Dataset (MSHD) is introduced as a comprehensive benchmark to systematically assess the robustness of existing methods on both common HOI and NCHOI in real-world applications. Extensive experimentation indicates that the proposed approach outperforms existing state-of-the-art techniques. It offers substantial improvements, particularly in NCHOI scenarios, which highlight its effectiveness in resolving issues related to long-range interactions.

Ruicheng Zhang, Haowei Guo, Kanghui Tian, Jun Zhou, Mingliang Yan, Zeyu Zhang 0006, Shen Zhao

Unified Medical Image Segmentation (UMIS) is critical for comprehensive anatomical assessment but faces challenges due to multi-scale structural heterogeneity. Conventional pixel-based approaches, lacking object-level anatomical insight and inter-organ relational modeling, struggle with morphological complexity and feature conflicts, limiting their efficacy in UMIS. We propose Mamba Snake, a novel deep snake framework enhanced by state space modeling for UMIS. Mamba Snake frames multi-contour evolution as a hierarchical state space atlas, effectively modeling macroscopic inter-organ topological relationships and microscopic contour refinements. We introduce a snake-specific vision state space module, the Mamba Evolution Block (MEB), which leverages effective spatiotemporal information aggregation for adaptive refinement of complex morphologies. Energy map shape priors further ensures robust long-range contour evolution in heterogeneous data. Additionally, a dual-classification synergy mechanism is incorporated to concurrently optimize detection and segmentation, mitigating under-segmentation of microstructures in UMIS. Extensive evaluations across five clinical datasets reveal Mamba Snake's superior performance.

Ruicheng Zhang, Yu Sun, Zeyu Zhang 0006, Jinai Li, Xiaofan Liu, Hoi Fan Au, Haowei Guo, Puxin Yan

We introduce MARL-MambaContour, the first contour-based medical image segmentation framework based on Multi-Agent Reinforcement Learning (MARL). Our approach reframes segmentation as a multi-agent cooperation task focused on generating topologically consistent object-level contours, addressing the limitations of traditional pixel-based methods which could lack topological constraints and holistic structural awareness of anatomical regions. Each contour point is modeled as an autonomous agent that iteratively adjusts its position to align precisely with the target boundary, enabling adaptation to blurred edges and intricate morphologies common in medical images. This iterative adjustment process is optimized by a contour-specific Soft Actor-Critic (SAC) algorithm, further enhanced with the Entropy Regularization Adjustment Mechanism (ERAM) which dynamically balances agent exploration with contour smoothness. Furthermore, the framework incorporates a Mamba-based policy network featuring a novel Bidirectional Cross-attention Hidden-state Fusion Mechanism (BCHFM). This mechanism mitigates potential memory confusion limitations associated with long-range modeling in state space models, thereby facilitating more accurate inter-agent information exchange and informed decision-making. Extensive experiments on five diverse medical imaging datasets demonstrate the state-of-the-art performance of MARL-MambaContour, highlighting its potential as an accurate and robust clinical application.

Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Haolan Zhan, Yixian Shen, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001 等

Underwater scenes present significant challenges for modern 3D scene reconstruction techniques due to absorption, in-scattering, and out-scattering effects, which alter light transport and degrade reconstruction quality, especially under sparse-view conditions. We present AtlantisGS, a novel underwater scene reconstruction method, which only requires sparse-view inputs. It incorporates a scattering decomposition method that separates medium and object contributions during rendering, and a sparse Gaussian proliferation strategy that adaptively densifies the scene representation to improve structural accuracy. These components jointly enhance both geometric reconstruction and medium modeling, enabling accurate and efficient scene recovery with limited observations. Extensive experiments on real-world underwater datasets demonstrate that AtlantisGS outperforms existing NeRF- and 3DGS-based methods across various metrics. AtlantisGS achieves higher reconstruction fidelity with significantly fewer input views and real-time rendering capability. These results establish AtlantisGS as an effective solution for sparse-view underwater 3D scene reconstruction.

Yunlong Zhao 0003, Xiaoheng Deng, Zhuohua Qiu, Feng Yang, Chang Xu 0002, Xiangjian He, Shan You, Xiu Su

Dynamic novel view synthesis (NVS) aims to render time-varying scenes from arbitrary viewpoints, balancing rendering quality and computational efficiency. While recent 4D Gaussian Splatting approaches offer promising real-time performance, they fundamentally overlook critical interdependence between Gaussians by modeling deformations independently. Our information-theoretic analysis reveals substantial mutual information across the Gaussian field, manifesting as appearance-preserving radiance coherence and motion-consistent deformation propagation. This finding establishes that rendering quality emerges from coordinated transformation rather than independent processing. We propose Correlation-aware Dynamic Gaussian Splatting (CaDGS) with our novel Gaussian Correlation Tensor Projection (GCTP) method, which efficiently transforms the complex O(n3) mutual information tensor into a dual-channel O(n2) spatial matrix, preserving the critical topological structure of Gaussian interactions. Combined with our Spatio-Temporal Deformation Consistency (STDC) learning, which enforces volumetric coherence through tensor-guided regularization across multiple scales, CaDGS prevents geometric distortions and texture inconsistencies common in previous approaches. Experimental results demonstrate state-of-the-art performance, achieving 32.4 PSNR on the Neu3D dataset with fewer Gaussians while maintaining rendering speeds of 323 FPS at 1353 × 1014 resolution.

Yong Liu 0031, Jinshan Pan, Yinchuan Li, Qingji Dong, Chao Zhu 0007, Yu Guo 0006, Fei Wang 0008

Diffusion models have shown great potential in generating realistic image detail. However, adapting these models to video super-resolution (VSR) remains challenging due to their inherent stochasticity and lack of temporal modeling. Previous methods have attempted to mitigate this issue by incorporating motion information and temporal layers. However, unreliable motion estimation from low-resolution videos and costly multiple sampling steps with deep temporal layers limit them to short sequences. In this paper, we propose UltraVSR, a novel framework that enables ultra-realistic and temporally-coherent VSR through an efficient one-step diffusion space. A central component of UltraVSR is the Degradation-aware Reconstruction Scheduling (DRS), which estimates a degradation factor from the low-resolution input and transforms the iterative denoising process into a single-step reconstruction from low-resolution to high-resolution videos. To ensure temporal consistency, we propose a lightweight Recurrent Temporal Shift (RTS) module, including an RTS-convolution unit and an RTS-attention unit. By partially shifting feature components along the temporal dimension, it enables effective propagation, fusion, and alignment across frames without explicit temporal layers. The RTS module is integrated into a pretrained text-to-image diffusion model and is further enhanced through Spatio-temporal Joint Distillation (SJD), which improves temporally coherence while preserving realistic details. Additionally, we introduce a Temporally Asynchronous Inference (TAI) strategy to capture long-range temporal dependencies under limited memory constraints. Extensive experiments show that UltraVSR achieves state-of-the-art performance, both qualitatively and quantitatively, in a single sampling step. Code is available at https://github.com/yongliuy/UltraVSR.

Sitian Gu, Zhiyu Pan, Chaoyi Hong, Chengxin Liu, Zhiguo Cao 0001

Video reframing, which converts landscape-oriented (LO) to portrait-oriented (PO) video for some PO devices such as smartphones and tablets, faces challenges. Existing approaches mainly follow a multi-step pipeline to preserve video content that ignore composition quality due to lack of large-scale datasets. To address these challenges, we propose a fully automated composition-aware dataset using vision-language models and image composition assessment models, pairing LO videos with high-quality PO versions. We then propose an end-to-end model with an attention-aware backbone and a time-aware consistency module. Experiments show our approach outperforms others in efficiency and effectiveness, proving that composition awareness and end-to-end modeling are critical for video reframing.

Kailong Yu, Liyuan Pan, Liu Liu 0009, Wei Liang 0008

Image de-reflection is a critical task in computer vision. Existing methods for de-reflection using monocular cameras face challenges due to the lack of depth cues to separate the transmission and reflection layers, particularly under strong illumination or multi-layer reflection scenarios. Although recent advances, such as 3D Gaussian Splatting (3DGS), utilize novel view-synthesis capabilities to separate transmitted and reflected layers, they still encounter difficulties in practice with monocular images. In this paper, we simplify the de-reflection task by combining dual-pixel (DP) technology with 3DGS, forming the first unsupervised de-reflection framework. Specifically, we propose the Dual-View Coordinated Reflection Removal (DCRR) Framework, which integrates depth cues from DP sensors with the rendering capabilities of 3DGS. The DCRR utilizes a dual-view approach that estimates the image transmission layer and opacity via differentiable rasterization with 3DGS and reconstructs the reflection layer through a lightweight multi-layer perceptron. We then present the Dual-Pixel-Driven Reflection Gaussian Pruning (DPRGP) to refine the separation process. By using the physical properties of DP sensors, DCRR achieves significant accuracy improvements in complex reflection scenarios. A real-world DP-based dataset that includes paired reflection/reflection-free images has been collected. Extensive experiments demonstrate our competitive performance compared to state-of-the-art de-reflection approaches.

Shuning Sun, Yu Zhang 0296, Chen Wu 0006, Dianjie Lu, Guijuan Zhang, Yang Wen, Zhuoran Zheng

Video imaging is often affected by complex degradations such as blur, noise, and compression artifacts. Traditional restoration methods follow a ''single-task single-model'' paradigm, resulting in poor generalization and high computational cost, limiting their applicability in real-world scenarios with diverse degradation types. We propose UniFlowRestore, a general video restoration framework that models restoration as a time-continuous evolution under a prompt-guided and physics-informed vector field. A physics-aware backbone PhysicsUNet encodes degradation priors as potential energy, while PromptGenerator produces task-relevant prompts as momentum. These components define a Hamiltonian system whose vector field integrates inertial dynamics, decaying physical gradients, and prompt-based guidance. The system is optimized via a fixed-step ODE solver to achieve efficient and unified restoration across tasks. Experiments show that UniFlowRestore delivers state-of-the-art performance with strong generalization and efficiency. Quantitative results demonstrate that UniFlowRestore achieves state-of-the-art performance, attaining the highest PSNR (33.89 dB) and SSIM (0.97) on the video denoising task, while maintaining top or second-best scores across all evaluated tasks.