论文检索

输入标题、作者或关键词,从 9,256 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
9,256篇论文匹配“Diffusion models”
第 174 / 463 页

Demin Yu, Wenchuan Du, Kenghong Lin, Xutao Li 0001, Yunming Ye, Chuyao Luo, Xunlai Chen

Precipitation nowcasting plays a pivotal role in urban planning and disaster mitigation, where extending forecast horizons offers critical advantages for proactive decision-making. Most data-driven methods focus on modeling radar echo sequences through end-to-end spatiotemporal predictive learning, yielding precise short-term predictions; however, they fundamentally neglect the inherent physical mechanism governing precipitation system. Moreover, approaches relying solely on single-modality radar observations suffer from persistent information bottlenecks, severely limiting their temporal generalizability for extended forecasting. To address these challenges, we propose PiMMNet, a Physics-informed Multi-Modal Network. It is constructed based on the advection-diffusion principle from fluid dynamics, explicitly modeling the precipitation evolution as a spatiotemporal transport processes characterized by the deterministic advection and the stochastic source. We carefully design a multi-model motion estimation network and a motion-guided diffusion model to describe the deterministic and stochastic terms, respectively. The core innovation of our method lies in jointly estimating a physics-constrained velocity field from multi-modal inputs (radar and satellite data). In this case, we naturally align the motion evolution among modalities into a unified representation, inherently mitigating cross-modal distribution biases. Experimental evaluations on two real-world multi-modal meteorological datasets demonstrate the efficacy of our approach, showcasing significant improvements in accuracy and robustness for longer-range precipitation nowcasting. Our code are available at https://github.com/DeminYu98/PiMMNet.

Yitong Sun 0002, Yao Huang, Ruochen Zhang, Huanran Chen, Shouwei Ruan, Ranjie Duan, Xingxing Wei 0001

Despite the impressive generative capabilities of text-to-image (T2I) diffusion models, they remain vulnerable to generating inappropriate content, especially when confronted with implicit sexual prompts. Unlike explicit harmful prompts, these subtle cues, often disguised as seemingly benign terms, can unexpectedly trigger sexual content due to underlying model biases, raising significant ethical concerns. However, existing detection methods are primarily designed to identify explicit sexual content and therefore struggle to detect these implicit cues. Fine-tuning approaches, while effective to some extent, risk degrading the model's generative quality, creating an undesirable trade-off. To address this, we propose NDM, the first noise-driven detection and mitigation framework, which could detect and mitigate implicit malicious intention in T2I generation while preserving the model's original generative capabilities. Specifically, we introduce two key innovations: first, we leverage the separability of early-stage predicted noise to develop a noise-based detection method that could identify malicious content with high accuracy and efficiency; second, we propose a noise-enhanced adaptive negative guidance mechanism that could optimize the initial noise by suppressing the prominent region's attention, thereby enhancing the effectiveness of adaptive negative guidance for sexual mitigation. Experimentally, we validate NDM on both natural and adversarial datasets, demonstrating its superior performance over existing SOTA methods, including SLD, UCE, and RECE, etc.

Junlei Zhou, Jiashi Gao, Xinwei Guo, Haiyan Wu, Quanying Liu, Xiangyu Zhao 0001, Hongxin Wei, Xin Yao 0001, Xuetao Wei

Text-to-Image (T2I) diffusion models exhibit concerning tendencies to generate harmful imagery that perpetuates social biases and stereotypes, posing significant ethical risks in real-world applications. While existing mitigation approaches predominantly employ black-box methodologies through dataset augmentation or constrained fine-tuning, they face critical limitations, including high data acquisition costs and potential exacerbation of stereotypes during model retraining. Inspired by neuroscience principles where neurological dysfunction often stems from aberrant neural activation patterns, we propose a novel framework, StereoClinic, targeting the root cause of stereotype generation through direct neural intervention. Our solution introduces two synergistic components: Diffusion Deep Taylor Decomposition (DDTD) for precisely localizing stereotype-related neurons via Layer-wise Relevance Propagation (LRP) attribution analysis, and Stereotype Neuron Suppression (SNS) implementing targeted activation damping to neutralize bias propagation. Through extensive empirical evaluations across multiple bias dimensions, we demonstrate that our method achieves significant stereotype mitigation without compromising image quality or requiring additional training data. This neuro-inspired approach establishes a new paradigm for model interpretability and ethical alignment in generative AI systems.

Poyuan Mao, Cheng-Chang Tsai, Chun-Shien Lu

The great success of the diffusion model in image synthesis led to the release of gigantic commercial models, raising the issue of copyright protection and inappropriate content generation. Training-free diffusion watermarking provides a low-cost solution for these issues. However, the prior works remain vulnerable to rotation, scaling, and translation (RST) attacks. Although some methods employ meticulously designed patterns to mitigate this issue, they often reduce watermark capacity, which can result in identity (ID) collusion. To address these problems, we propose MaXsive, a training-free diffusion model generative watermarking technique that has high capacity and robustness. MaXsive best utilizes the initial noise to watermark the diffusion model. Moreover, instead of using a meticulously repetitive ring pattern, we propose injecting the X-shape template to recover the RST distortions. This design significantly increases robustness without losing any capacity, making ID collusion less likely to happen. The effectiveness of MaXsive has been verified on two well-known watermarking benchmarks under the scenarios of verification and identification.

Shengjiu Dai, Xiujian Liang, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001

Bitwise Vision AutoRegressive (BVAR) Model, as a distinguished source of young blood, has been taking the lead in the track of text-to-image synthesis, which at the same time raises legal and ethnic concerns such as copyright and authenticity. However, existing methods mainly focus on watermarking within diffusion models, which rely on the distinctive attributes of diffusion steps and cannot be directly transferred to new circumstances. To this end, we propose Safe-BVAR, the first watermark framework to embed bit strings during image generation in BVAR. Our study discovers the local similarity of the inferenced latent feature and the element-wise robustness of image autoencoder. Therefore, combined with the residual-accumulative nature of BVAR, we propose a novel Late Stage Residual Implanter to embed watermark and extract the information based on Local Contextual Extractor. Furthermore, we propose a Distributed Rotational Arranger to enhance watermark against local distortions. Our method is training-free and plug-and-play. Meanwhile, it can be easily applied to flexible-sized images. We evaluate the robustness and invisibility of the watermark, showing that it can resist common image attacks and cast inappreciable influence on the image.

Naresh Kumar Devulapally, Shruti Agarwal, Tejas Gokhale, Vishnu Suresh Lokhande

Text-to-image diffusion models have demonstrated remarkable effectiveness in rapid and high-fidelity personalization, even when provided with only a few user images. However, the effectiveness of personalization techniques has lead to concerns regarding data privacy, intellectual property protection, and unauthorized usage. To mitigate such unauthorized usage and model replication, the idea of generating ''unlearnable'' training samples utilizing image poisoning techniques has emerged. Existing methods for this have limited imperceptibility as they operate in the pixel space which results in images with noise and artifacts. In this work, we propose a novel model-based perturbation strategy that operates within the latent space of diffusion models. Our method alternates between denoising and inversion while modifying the starting point of the denoising trajectory: of diffusion models. This trajectory-shifted sampling ensures that the perturbed images maintain high visual fidelity to the original inputs while being resistant to inversion and personalization by downstream generative models. This approach integrates unlearnability into the framework of Latent Diffusion Models (LDMs), enabling a practical and imperceptible defense against unauthorized model adaptation. We validate our approach on four benchmark datasets to demonstrate robustness against state-of-the-art inversion attacks. Results demonstrate that our method achieves significant improvements in imperceptibility (~8% - 10% on perceptual metrics including PSNR, SSIM, and FID) and robustness (~10% on average across five adversarial settings), highlighting its effectiveness in safeguarding sensitive data. https://github.com/naresh-ub/unlearnable_samples.

Shunchang Liu, Zhuan Shi, Lingjuan Lyu, Yaochu Jin, Boi Faltings

Assessing whether AI-generated images are substantially similar to copyrighted works is a crucial step in resolving copyright disputes. In this paper, we propose CopyJudge, an automated copyright infringement identification framework that leverages large vision-language models (LVLMs) to simulate practical court processes for determining substantial similarity between copyrighted images and those generated by text-to-image diffusion models. Specifically, we employ an abstraction-filtration-comparison test framework with multi-LVLM debate to assess the likelihood of infringement and provide detailed judgment rationales. Based on the judgments, we further introduce a general LVLM-based mitigation strategy that automatically optimizes infringing prompts by avoiding sensitive expressions while preserving the non-infringing content. Besides, our approach can be enhanced by exploring non-infringing noise vectors within the diffusion latent space via reinforcement learning, even without modifying the original prompts.Experimental results show that our identification method achieves comparable state-of-the-art performance, while offering superior generalization and interpretability across various forms of infringement, and that our mitigation method could more effectively mitigate memorization and IP infringement without losing non-infringing expressions.

Yizhou Lin, Nisha Huang, Kaer Huang, Henglin Liu, Yiqiang Yan, Jie Guo, Tong-Yee Lee, Xiu Li 0001

The success of diffusion models in text-to-image (T2I) generation has made it urgent to remove unwanted concepts, such as copyrighted, offensive, and unsafe ones, from pre-trained models in an accurate, timely, and cost-effective manner. However, limited by the inherent optimization perspective, existing methods have two major problems. Firstly, they overlook maintaining the global visual style during the erasure process, leading to significant style shifts. Secondly, excessive concept erasure causes relevant content to disappear or generates substitutes unrelated to the original object's attributes. Compared to other methods, our proposed ICE has unique advantages, as it can generate diverse visual features and achieve a balance between concept erasure and maintaining the semantic content of the target object. This is mainly achieved through our well-designed non-erasable features protector (NEFP) and augmented invariant constraints (AIC). Specifically, we enhance the protection of feature information by embedding an augmented orthogonal anchor concept matrix. Meanwhile, under controlled constraints, we introduce invariants into the embedding space to retain key semantics. This work specifically emphasizes the importance of focusing on feature expression and semantic protection in the concept erasure task for fully unleashing the performance of T2I models.

Fan Qi, Ao Liu, Zixin Zhang 0004, Changsheng Xu

Federated face generation technology leverages decentralized private data to achieve high-quality face synthesis. However, regulations such as the GDPR confer users the right to be forgotten, necessitating the removal of contributions from specific clients in the global model. Existing generation model unlearning methods are primarily designed for centralized environments and are inadequate for addressing the constraints of data privacy storage and limited client computational resources in federated settings. To address this gap, we propose F2GU, the first federated unlearning framework specifically tailored for face generation models, enabling the effective removal of contributions associated with specific clients (identities) while ensuring privacy. Our proposed Generation Trajectory Redirection method dynamically guides the generation trajectory away from target identities, thereby effectively eliminating contributions from specific clients. Additionally, we devise a Mirroring-guided Trajectory Optimization strategy that constructs a mirror projection utilizing the retained client trajectory origins to ensure the generative capabilities of the model are preserved post-unlearning. We conduct extensive experiments on two mainstream face generation models (GAN and Diffusion Model) across three different datasets.The results indicate that our method demonstrates superior performance in both the success rate of identity unlearning and the preservation of generation quality. The code can be available at https://github.com/FanQi-AI/FFGU.

Song Yan 0001, Hui Wei 0004, Jinlong Fei, Guoliang Yang 0005, Zhengyu Zhao 0001, Zheng Wang 0007

Various (text) prompt filters and (image) safety checkers have been implemented to mitigate the misuse of Text-to-Image (T2I) models in creating Not-Safe-For-Work (NSFW) content. In order to expose potential security vulnerabilities of such safeguards, multimodal jailbreaks have been studied. However, existing jailbreaks are limited to prompt-specific and image-specific perturbations, which suffer from poor scalability and time-consuming optimization. To address these limitations, we propose Universally Unfiltered and Unseen (U3)-Attack, a multimodal jailbreak attack method against T2I safeguards. Specifically, U3-Attack optimizes an adversarial patch on the image background to universally bypass safety checkers and optimizes a safe paraphrase set from a sensitive word to universally bypass prompt filters while eliminating redundant computations. Extensive experimental results demonstrate the superiority of our U3-Attack on both open-source and commercial T2I models. For example, on the commercial Runway-inpainting model with both prompt filter and safety checker, our U3-Attack achieves approximately 4× higher success rates than the state-of-the-art multimodal jailbreak attack, MMA-Diffusion. Content Warning: This paper includes examples of NSFW content.

Yujiang Li, Zhili Zhou 0001, Ruohan Meng, Baowei Wang, Xiaojuan Wang, Cheng Qiao, Jiantao Zhou 0001

While the widespread adoption of diffusion models in image generation has showcased remarkable capabilities, it has also inadvertently opened the door to malicious exploitation. Recent research has primarily concentrated on protecting images from the misuse of diffusion-based customized generation (CG). However, these approaches often overlook that image details can still be enhanced through diffusion-based super-resolution (SR) techniques, significantly increasing the risks of personal image leakage and abuse. To combat these multifaceted risks, we propose the Zero Matrix-guided Adaptive Image Vaccine (ZMAIV) framework. Specifically, we introduce the Self-attention Removal strategy, tailored for CG, which disrupts the model's core mechanism of focusing on sensitive spaces. Concurrently, the High-frequency Removal strategy is proposed to impede the high-frequency details reconstruction of SR. These defense strategies effectively dismantle the underlying mechanisms that facilitate unauthorized data extrapolation. Moreover, the proposed Adaptive Space Search Attack precisely targets critical spaces within images for vaccine injection, optimizing perturbation placement to minimize perturbation conflict while maintaining defense performance. Extensive experiments demonstrate that the proposed ZMAIV outperforms the state-of-the-arts in the aspects of simultaneously defending against diffusion-based CG and SR, affirming its superiority in safeguarding visual content against these dual threats.

Yilin Lu, Jianghang Lin, Linhuang Xie, Kai Zhao 0013, Yansong Qu, Shengchuan Zhang, Liujuan Cao, Rongrong Ji

Anomaly inspection plays a vital role in industrial manufacturing, but the scarcity of anomaly samples significantly limits the effectiveness of existing methods in tasks such as localization and classification. While several anomaly synthesis approaches have been introduced for data augmentation, they often struggle with low realism, inaccurate mask alignment, and poor generalization. To overcome these limitations, we propose Generate Aligned Anomaly (GAA), a region-guided, few-shot anomaly image-mask pair generation framework. GAA leverages the strong priors of a pretrained latent diffusion model to generate realistic, diverse, and semantically aligned anomalies using only a small number of samples. The framework first employs Localized Concept Decomposition to jointly model the semantic features and spatial information of anomalies, enabling flexible control over the type and location of anomalies. It then utilizes Adaptive Multi-Round Anomaly Clustering to perform fine-grained semantic clustering of anomaly concepts, thereby enhancing the consistency of anomaly representations. Subsequently, a region-guided mask generation strategy ensures precise alignment between anomalies and their corresponding masks, while a low-quality sample filtering module is introduced to further improve the overall quality of the generated samples. Extensive experiments on the MVTec AD and LOCO datasets demonstrate that GAA achieves superior performance in both anomaly synthesis quality and downstream tasks such as localization and classification.

Liu Yu 0001, Jiajun Sun, Ping Kuang, Rui Zhou 0012, Fan Zhou 0002, Zhikun Feng

Social biases in text-to-image models have drawn increasing attention, yet existing debiasing efforts often focus solely on either the textual (e.g., CLIP) or visual (e.g., U-Net) space. This unimodal perspective introduces two major challenges: (i) Debiasing only the textual space fails to control visual outputs, often leading to pseudo- or over-corrections due to unaddressed visual biases during denoising; (ii) Debiasing only the visual space can cause modality conflicts when biases in textual and vision are misaligned, degrading the quality and consistency of generated images. To address these issues, we propose a Bimodal ADaptive Guidance DEbiasing within Textual and Visual Spaces (BADGE). First, BADGE quantifies attribute-level bias inclination in both modalities, providing precise guidance for subsequent mitigation. Second, to avoid pseudo/over-correction and modality conflicts, the quantified bias degree is used as the debiasing strength for adaptive guidance, enabling fine-grained correction tailored to discrete attribute concepts.Extensive experiments demonstrate that BADGE significantly enhances fairness across intra- and inter-category attributes (e.g., gender, skin tone, age, and their interaction) while preserving high image fidelity. *Our project page is at https://badgediffusion.github.io/

Jiadong Pan, Liang Li 0003, Hongcheng Gao, Zheng-Jun Zha, Qingming Huang, Jiebo Luo 0001

Diffusion models (DMs) have demonstrated exceptional performance in text-to-image tasks, leading to their widespread use. With the introduction of classifier-free guidance (CFG), the quality of images generated by DMs is significantly improved. However, one can use DMs to generate more harmful images by maliciously guiding the image generation process through CFG. Existing safe alignment methods aim to mitigate the risk of generating harmful images but often reduce the quality of clean image generation. To address this issue, we propose SafeCFG to adaptively control harmful features with dynamic safe guidance by modulating the CFG generation process. It dynamically guides the CFG generation process based on the harmfulness of the prompts, inducing significant deviations only in harmful CFG generations, achieving high quality and safety generation. SafeCFG can simultaneously modulate different harmful CFG generation processes, so it could eliminate harmful elements while preserving high-quality generation. Additionally, SafeCFG provides the ability to detect image harmfulness, allowing unsupervised safe alignment on DMs without pre-defined clean or harmful labels. Experimental results show that images generated by SafeCFG achieve both high quality and safety, and safe DMs trained in our unsupervised manner also exhibit good safety performance. The project page is https://github.com/matrix0721/SafeCFG.

Jiaqi Xu, Kunzhe Huang, Xinyi Zou, Yunkuo Chen, Bo Liu, Mengli Cheng, Jun Huang 0007, Xing Shi

This paper introduces EasyAnimate, an efficient and high quality video generation framework that leverages diffusion transformers to achieve high-quality video production, encompassing data processing, model training, and end-to-end inference. Despite substantial advancements achieved by video diffusion models, existing video generation models still struggles with slow generation speeds and less-than-ideal video quality. To improve training and inference efficiency without compromising performance, we propose Hybrid Window Attention. We design the multidirectional sliding window attention in Hybrid Window Attention, which provides stronger receptive capabilities in 3D dimensions compared to naive one, while reducing the model's computational complexity as the video sequence length increases. To enhance video generation quality, we optimize EasyAnimate using reward backpropagation to better align with human preferences. As a post-training method, it greatly enhances the model's performance while ensuring efficiency. In addition to the aforementioned improvements, EasyAnimate integrates a series of further refinements that significantly improve both computational efficiency and model performance. We introduce a new training strategy called Training with Token Length to resolve uneven GPU utilization in training videos of varying resolutions and lengths, thereby enhancing efficiency. Additionally, we use a multimodal large language model as the text encoder to improve text comprehension of the model. Experiments demonstrate significant enhancements resulting from the above improvements. The EasyAnimate achieves state-of-the-art performance on both the VBench leaderboard and human evaluation. Code and pre-trained models are available at https://github.com/aigc-apps/EasyAnimate.

Chenxi Li, Weijie Wang 0002, Qiang Li 0048, Nicu Sebe, Bruno Lepri, Weizhi Nie

Text-driven object insertion in the 3D scene is an emerging task that enables intuitive scene editing through natural language. Despite its potential, existing 2D editing-based methods often suffer from reliance on spatial priors such as 2D masks, 3D bounding boxes, and they struggle to ensure inserted object consistency. These limitations hinder flexibility and scalability in real-world applications. In this paper, we propose FreeInsert, a novel framework that leverages foundation models (MLLMs, LGM, and diffusion models) to disentangle object generation and spatial placement, enabling unsupervised and flexible object insertion in 3D scenes without spatial priors. FreeInsert begins with an MLLM-based parser that extracts structured semantics-including object types, spatial relationships, and attachment regions-from user instructions. These semantics guide both the reconstruction of the inserted object for 3D consistency and the learning of its degrees of freedom. We first leverage the spatial reasoning capabilities of MLLMs to initialize the object's pose and scale. To further enhance natural integration with the scene, a hierarchical spatially-aware stage is employed to refine the object's placement, incorporating both the spatial semantics and priors inferred by the MLLM. Finally, the object's appearance is enhanced using inserted-object image to improve visual fidelity. Experimental results demonstrate that FreeInsert enables semantically coherent, spatially precise, and visually realistic 3D insertions, without requiring any spatial priors, offering a user-friendly and flexible editing experience. Project page: https://tjulcx.github.io/FreeInsert/.

Chunshi Wang, Hongxing Li, Yawei Luo

While 3D Gaussian representations (3DGS) have proven effective for modeling the geometry and appearance of objects, their potential for capturing other physical attributes-such as sound-remains largely unexplored. In this paper, we present a novel framework dubbed SonicGauss for synthesizing impact sounds from 3DGS representations by leveraging their inherent geometric and material properties. Specifically, we integrate a diffusion-based sound synthesis model with a PointTransformer-based feature extractor to infer material characteristics and spatial-acoustic correlations directly from Gaussian ellipsoids. Our approach supports spatially varying sound responses conditioned on impact locations and generalizes across a wide range of object categories. Experiments on the ObjectFolder dataset and real-world recordings demonstrate that our method produces realistic, position-aware auditory feedback. The results highlight the framework's robustness and generalization ability, offering a promising step toward bridging 3D visual representations and interactive sound synthesis.

Hongjie Wu, Mingqin Zhang, Linchao He, Ji-Zhe Zhou 0001, Jiancheng Lv 0001

Diffusion models have shown remarkable promise for image restoration by leveraging powerful priors. Prominent methods typically frame the restoration problem within a Bayesian inference framework, which iteratively combines a denoising step with a likelihood guidance step. However, the interactions between these two components in the generation process remain underexplored. In this paper, we analyze the underlying gradient dynamics of these components and identify significant instabilities. Specifically, we demonstrate conflicts between the prior and likelihood gradient directions, alongside temporal fluctuations in the likelihood gradient itself. We show that these instabilities disrupt the generative process and compromise restoration performance. To address these issues, we propose Stabilized Progressive Gradient Diffusion (SPGD), a novel gradient management technique. SPGD integrates two synergistic components: (1) a progressive likelihood warm-up strategy to mitigate gradient conflicts; and (2) adaptive directional momentum (ADM) smoothing to reduce fluctuations in the likelihood gradient. Extensive experiments across diverse restoration tasks demonstrate that SPGD significantly enhances generation stability, leading to state-of-the-art performance in quantitative metrics and visually superior results. Code is available at https://github.com/74587887/SPGD.

Jeongsoo Choi, Ji-Hoon Kim, Sung-Bin Kim, Tae-Hyun Oh, Joon Son Chung

In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attention due to its wide range of applications, such as film production, dubbing, and virtual avatars. Despite recent progress, existing methods still suffer from limitations in speech intelligibility, audio-video synchronization, speech naturalness, and voice similarity to the reference speaker. To address these challenges, we propose AlignDiT, a multimodal Aligned Diffusion Transformer that generates accurate, synchronized, and natural-sounding speech from aligned multimodal inputs. Built upon the in-context learning capability of the DiT architecture, AlignDiT explores three effective strategies to align multimodal representations. Furthermore, we introduce a novel multimodal classifier-free guidance mechanism that allows the model to adaptively balance information from each modality during speech synthesis. Extensive experiments demonstrate that AlignDiT significantly outperforms existing methods across multiple benchmarks in terms of quality, synchronization, and speaker similarity. Moreover, AlignDiT exhibits strong generalization capability across various multimodal tasks, such as video-to-speech synthesis and visual forced alignment, consistently achieving state-of-the-art performance. The demo page is available at https://mm.kaist.ac.kr/projects/AlignDiT.

Yuntian Xiao, Shoulong Zhang, Zihang Zhang, Jiahao Cui 0001, Yan Wang, Shuai Li 0001

Generating highly realistic 4D interaction in real time is significant for visual content generation. Although existing works have validated to produce impressive dynamics by employing physical simulation and learned material mainly from pre-trained video diffusion models, it is still challenging to generate real-time 4D interaction with high-quality motion due to the heavy time consumption of the simulation solver and indirect material learning strategy. This paper proposes a novel physics-based 4D generation method, Phys4DRT, for arbitrary realistic real-time interaction on 3D Gaussian Splatting (3DGS) objects with direct motion supervision in time-frequency domain. Specifically, we devise a fast and differentiable eXtended Position Based Dynamics (XPBD) simulator as the light-weight controller for efficient physical evolution on a quasi-regular tetrahedral proxy mesh, into which we immerse the static 3DGS for efficient and stable deformation simulation. In addition, to learn the heterogeneous material for realistic motion, we directly supervise the generated dynamic 3D behavior by the motion representation of the optical flow and spectral volume extracted from the generated reference video, rather than indirect supervision in the color space used in previous approaches. We thoroughly conduct experiments on the public benchmarks to demonstrate the efficiency and effectiveness of our method. Our model can accelerate real-time 4D interaction generation by approximately x20 faster than the current Material Point Method (MPM) based approaches while achieving competitive visual quality compared with the state-of-the-art baselines.