论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 22 / 81 页

Bingqian Zhou, Zhihao Wu, Yushi Cheng, Wenyuan Xu 0001

Text-guided inpainting models are widely used for image editing, restoration, and content generation due to their ability to produce high-fidelity results aligned with natural language prompts. However, these models remain vulnerable to jailbreaking attacks, where adversaries manipulate inputs to generate pornographic or violent content. While prior attacks rely on adversarial text prompts, they are increasingly mitigated by advanced text-based safety filters and manual review. In this work, we propose a new attack paradigm that bypasses these defenses by leveraging the image modality alone. Specifically, we inject imperceptible adversarial perturbations into the input image, enabling successful jailbreaks even when paired with clean prompts (e.g., ''a woman''). To achieve this, we address two key challenges: (1) stabilizing the optimization of adversarial perturbations via a novel gradient estimator, and (2) ensuring visual imperceptibility through a diffusion-based perturbation generator. Extensive experiments show that our method successfully compromises the Stable Diffusion Inpainting model-despite its built-in image and text safety checkers-achieving an average attack success rate (ASR) of 85.7%, significantly outperforming baselines (58.7%). Moreover, our attack exhibits strong transferability across models and maintains robustness against common image pre-processing defenses. Warning: Blurred or masked NSFW imagery is contained.

Xuan Hai, Xin Liu 0050, Zihao Zhang, Ziyao Yu, Xiangzhen Kong, Song Li 0006, Weina Niu, Rui Zhou 0005, Qingguo Zhou

The application of deep learning in voice cloning has significantly enhanced the quality of cloned voices. While advanced voice cloning technologies are widely applied across various domains, they also pose serious security challenges such as producing natural Deepfakes. In response, numerous studies have focused on detecting fake voices, with many reporting outstanding performance. However, is the issue truly resolved? This paper introduces Adversarial Neural Mimicry Attack (ANMA) which leverages a specialized model to predict the behavior of other similar models, transforming black-box attacks into white-box scenarios indirectly. Based on ANMA and Speaker-irrelative Features (SiFs), we propose a novel black-box attack framework called SiFMimicEvader, designed to evade fake voice detectors with high success rates and minimal query requirements. The framework utilizes speech representation models as the breakthrough to predict the behaviors of fake voice detectors and employs a series of SiFs editing operations as perturbations to deceive these detectors. Experimental results demonstrate the effectiveness of SiFMimicEvader, achieving an average attack success rate exceeding 50% across various detectors, significantly outperforming other attack methods, while also showing great performance in audio quality and query scale, indicating its high availability in real-world scenarios.

Jianqiao Cui, Bingyao Yu, Qihao Wang, Fei Meng, Jiwen Lu

This paper addresses the critical challenge of detecting codec-based audio deepfakes in multilingual and dynamically evolving adversarial scenarios. While existing detection systems exhibit performance degradation against codec-generated forgeries and unseen linguistic environments, we propose a novel audio deepfake detection framework ''WhiADD'' enhanced by semantic-acoustic fusion and cross-modal generalization. Our methodology introduces three key innovations: (1) The Union CodecFake (UCF) dataset, synthesized by extending the CodecFake generation pipeline to the multilingual Common Voice corpus, significantly expands acoustic diversity with 1.9M samples across varied phonetic, channel, and codec manipulation patterns. (2) A semantic-prompted Whisper architecture that integrates full-transcript linguistic constraints into decoder fine-tuning, enabling detection of semantic inconsistencies. (3) A gated cross-attention mechanism that dynamically fuses multi-source audio features with the proposed model's frozen encoder outputs, enhancing artifact detection through adaptive attention to pre-trained representations. Extensive experiments demonstrate state-of-the-art performance, achieving 0.55% EER on UCF testing data and less than 3% EER in zero-shot cross-lingual detection (German, French, Italian). The framework reduces false negatives by up to 24% compared to conventional models through improved semantic-acoustic alignment. These advancements establish a robust paradigm for combating evolving codec-based forgeries, bridging the critical gap between acoustic feature engineering and semantic coherence analysis in audio forensics.

Mingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang 0001, Haolin He, Zhiming Wang, Huijia Zhu, Jian Liu, Weiqiang Wang 0002

With the rapid advancement of AIGC technologies, audio deepfakes have become increasingly realistic, posing serious threats to information security and biometric authentication. Therefore, audio deepfake detection (ADD) has emerged as a critical and fast-evolving research area, particularly requiring superior generalization in out-of-domain scenarios. However, existing ADD methods suffer from constrained generalization and limited access to target data. To address these challenges, we propose Risk-Aware Style Alignment (RASA), a novel generalizable ADD framework that projects the style of any input feature into a shared style space through similarity-based projection. This alignment reduces both inter-domain and intra-source discrepancies without requiring target data during training. In addition, we adopt Structural Empirical Risk Minimization (SERM) in the Poincaré ball model to capture the hierarchical structure of the data and further minimize source risk. By jointly optimizing RASA and SERM, the proposed method effectively tightens the theoretical upper bound of target risk across three key dimensions: source risk, inter-domain divergence, and intra-source discrepancy. Extensive experiments demonstrate that our approach achieves superior generalization and outperforms existing state-of-the-art methods.

Cong Kong, Rui Xu, Jiawei Chen, Zhaoxia Yin

With the advancement of intelligent healthcare, medical pre-trained language models (Med-PLMs) have emerged and demonstrated significant effectiveness in downstream medical tasks. While these models are valuable assets, they are vulnerable to misuse and theft, requiring copyright protection. However, existing watermarking methods for pre-trained language models (PLMs) cannot be directly applied to Med-PLMs due to domain-task mismatch and inefficient watermark embedding. To fill this gap, we propose the first training-free backdoor model watermarking for Med-PLMs, employing low-frequency words as triggers and embedding the watermark by replacing their embeddings in the model's word embedding layer with those of specific medical terms. The watermarked Med-PLMs produce the same output for triggers as for the corresponding specified medical terms. We leverage this unique mapping to design tailored watermark extraction schemes for different downstream tasks, addressing the challenge of domain-task mismatch in previous methods. Experiments demonstrate superior effectiveness of our watermarking method across medical downstream tasks, robustness against model extraction, pruning, fusion-based backdoor removal attacks, and high efficiency with 10-second embedding. Our code is available at https://github.com/edu-yinzhaoxia/Med-PLMW.

Wenbo Xu, Junyan Wu, Wei Lu 0001, Xiangyang Luo 0001, Qian Wang 0002

Current researches on Deepfake forensics often treat detection as a classification task or temporal forgery localization problem, which are usually restrictive, time-consuming, and challenging to scale for large datasets. To resolve these issues, we present a multimodal deviation perceiving framework for weakly-supervised temporal forgery localization (MDP), which aims to identify temporal partial forged segments using only video-level annotations. The MDP proposes a novel multimodal interaction mechanism (MI) and an extensible deviation perceiving loss to perceive multimodal deviation, which achieves the refined start and end timestamps localization of forged segments. Specifically, MI introduces a temporal property preserving cross-modal attention to measure the relevance between the visual and audio modalities in the probabilistic embedding space. It could identify the inter-modality deviation and construct comprehensive video features for temporal forgery localization. To explore further temporal deviation for weakly-supervised learning, an extensible deviation perceiving loss has been proposed, aiming at enlarging the deviation of adjacent segments of the forged samples and reducing that of genuine samples. Extensive experiments demonstrate the effectiveness of the proposed framework and achieve comparable results to fully-supervised approaches in several evaluation metrics.

Midou Guo, Qilin Yin, Wei Lu 0001, Xiangyang Luo 0001

With the development of generative artificial intelligence, new forgery methods are rapidly emerging. Social platforms are flooded with vast amounts of unlabeled synthetic data and authentic data, making it increasingly challenging to distinguish real from fake. Due to the lack of labels, existing supervised detection methods struggle to effectively address the detection of unknown deepfake methods. Moreover, in open world scenarios, the amount of unlabeled data greatly exceeds that of labeled data. Therefore, we define a new deepfake detection generalization task which focuses on how to achieve efficient detection of large amounts of unlabeled data based on limited labeled data to simulate a open world scenario. To solve the above mentioned task, we propose a novel Open-World Deepfake Detection Generalization Enhancement Training Strategy (OWG-DS) to improve the generalization ability of existing methods. Our approach aims to transfer deepfake detection knowledge from a small amount of labeled source domain data to large-scale unlabeled target domain data. Specifically, we introduce the Domain Distance Optimization (DDO) module to align different domain features by optimizing both inter-domain and intra-domain distances. Additionally, the Similarity-based Class Boundary Separation (SCBS) module is used to enhance the aggregation of similar samples to ensure clearer class boundaries, while an adversarial training mechanism is adopted to learn the domain-invariant features. Extensive experiments show that the proposed deepfake detection generalization enhancement training strategy excels in cross-method and cross-dataset scenarios, improving the model's generalization.

Liqin Wang, Qianyue Hu, Wei Lu 0001, Xiangyang Luo 0001

The success of face recognition (FR) systems has led to serious privacy concerns due to potential unauthorized surveillance and user tracking on social networks. Existing methods for enhancing privacy fail to generate natural face images that can protect facial privacy. In this paper, we propose diffusion-based adversarial identity manipulation (DiffAIM) to generate natural and highly transferable adversarial faces against malicious FR systems. To be specific, we manipulate facial identity within the low-dimensional latent space of a diffusion model. This involves iteratively injecting gradient-based adversarial identity guidance during the reverse diffusion process, progressively steering the generation toward the desired adversarial faces. The guidance is optimized for identity convergence towards a target while promoting semantic divergence from the source, facilitating effective impersonation while maintaining visual naturalness. We further incorporate structure-preserving regularization to preserve facial structure consistency during manipulation. Extensive experiments on both face verification and identification tasks demonstrate that compared with the state-of-the-art, DiffAIM achieves stronger black-box attack transferability while maintaining superior visual quality. We also demonstrate the effectiveness of the proposed approach for commercial FR APIs, including Face++ and Aliyun.

Xinyu Xia 0002, Xingjun Ma, Yunfeng Hu 0003, Ting Qu 0001, Hong Chen 0003, Xun Gong 0007

Ensuring robust and generalizable autonomous driving requires not only broad scenario coverage but also efficient repair of failure cases, particularly those related to challenging and safety-critical scenarios. However, existing scenario generation and selection methods often lack adaptivity and semantic relevance, limiting their impact on performance improvement. In this paper, we propose SERA, an LLM-powered framework that enables autonomous driving systems to self-evolve by repairing failure cases through targeted scenario recommendation. By analyzing performance logs, SERA identifies failure patterns and dynamically retrieves semantically aligned scenarios from a structured bank. An LLM-based reflection mechanism further refines these recommendations to maximize relevance and diversity. The selected scenarios are used for few-shot fine-tuning, enabling targeted adaptation with minimal data. Experiments on the benchmark show that SERA consistently improves key metrics across multiple autonomous driving baselines, demonstrating its effectiveness and generalizability under safety-critical conditions.

Yan Wang 0088, Qindong Sun, Dongzhu Rong

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has enabled deepfake videos to evolve from unimodal generation to audio-visual forgeries. Existing multimodal deepfake detection methods primarily rely on capturing correlations between audio-visual modalities to improve detection performance. However, in real-world scenarios, network jitter often leads to audio-visual asynchrony, disrupting inter-modal associations and limiting the effectiveness of these methods. To address this issue, we propose a deepfake detection method specifically designed for audio-visual asynchrony scenarios. First, based on the theory of open balls in metric space, we analyze the variation mechanism of joint features in both audio-visual synchrony and asynchrony scenarios, revealing the impact of audio-visual asynchrony on detection performance. Second, we design a multimodal subspace representation module to mitigate inconsistencies in feature distributions and representation heterogeneity between modalities. We then formulate audio-visual feature alignment as an integer linear programming task and employ the Hungarian algorithm to reconstruct missing inter-modal associations. Finally, we introduce a self-supervised masked reconstruction mechanism to reconstruct missing features and construct the joint correlation matrix to measure cross-modal dependencies, enhancing the robustness of detection. Extensive experiments demonstrate that our method outperforms baselines in audio-visual asynchrony scenarios and exhibits robustness against unknown disturbances.

Jilong Wei, Yangyang Hu, Xiangjuan Wu, Yiqiang Wu, Hao Liu 0019

In this paper, we propose a Dual-Constraint Diffusion Model (DCDM) to contrast aging appearance for facial age estimation, addressing the key issue of noisy labels. Existing methods for face age estimation are plagued by class imbalance and noisy supervision signals, which disrupt the ordinal relationships between age categories and hinder effective feature decoupling in existing models. To overcome these challenges, the proposed DCDM develops a label-independent Paired Comparison, ensuring accurate sample labeling and maintaining continuity in age estimation. Moreover, we incorporate a Dual-Constraint Diffusion Model to effectively separate and recombine age-related and unrelated features, thus facilitating the generation of high-fidelity and continuous age-progressed facial representations. Lastly, we optimize our model parameters by exploiting the age difference information via an active learning framework. Comparative evaluations on several in-the-wild datasets demonstrate that our DCDM significantly achieves superior results compared to existing state-of-the-art methods in facial age estimation.

Demin Yu, Wenchuan Du, Kenghong Lin, Xutao Li 0001, Yunming Ye, Chuyao Luo, Xunlai Chen

Precipitation nowcasting plays a pivotal role in urban planning and disaster mitigation, where extending forecast horizons offers critical advantages for proactive decision-making. Most data-driven methods focus on modeling radar echo sequences through end-to-end spatiotemporal predictive learning, yielding precise short-term predictions; however, they fundamentally neglect the inherent physical mechanism governing precipitation system. Moreover, approaches relying solely on single-modality radar observations suffer from persistent information bottlenecks, severely limiting their temporal generalizability for extended forecasting. To address these challenges, we propose PiMMNet, a Physics-informed Multi-Modal Network. It is constructed based on the advection-diffusion principle from fluid dynamics, explicitly modeling the precipitation evolution as a spatiotemporal transport processes characterized by the deterministic advection and the stochastic source. We carefully design a multi-model motion estimation network and a motion-guided diffusion model to describe the deterministic and stochastic terms, respectively. The core innovation of our method lies in jointly estimating a physics-constrained velocity field from multi-modal inputs (radar and satellite data). In this case, we naturally align the motion evolution among modalities into a unified representation, inherently mitigating cross-modal distribution biases. Experimental evaluations on two real-world multi-modal meteorological datasets demonstrate the efficacy of our approach, showcasing significant improvements in accuracy and robustness for longer-range precipitation nowcasting. Our code are available at https://github.com/DeminYu98/PiMMNet.

Wenpeng Mu, Zheng Li, Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun

As image-generative AI models become increasingly accessible to the public, the demand for content safety has surged. Although model developers have introduced alignment mechanisms to prevent the creation of threatening images, and extensive researches have been conducted on verifying the authenticity of AI-generated images, a significant number of ex-regulatory images have been discovered that fall into regulatory gaps. These images are neither covered by existing alignment mechanisms nor included in the scope of current detection methods. To address this, we introduce ExDA, a detection and attribution framework specifically designed for such ex-regulatory images. ExDA utilizes a frozen CLIP:ViT-L/14 as a visual feature extractor to extract rich and unbiased visual features, complemented by a text feature reduction layer to unify semantic styles. For obtaining highly discriminative features, ExDA introduces an SFS-ResNet network, where each basic layer is replaced with a meticulously designed Multi-Channel Margin Convolution (MMConv). Additionally, a plug-and-play multi-generation model attributor is integrated behind the detector. Given the lack of ex-regulatory images in existing public datasets, we constructed ExImage, a dataset containing 72,000 ex-regulatory images, to validate ExDA's effectiveness. Experiments show that ExDA achieves an average detection accuracy of 99.07% on ExImage, and demonstrating significant performance improvements of +5.73% and +10.36% on GenImage and high-challenge Chameleon datasets respectively in cross-datasets evaluation. Notably, ExDA also achieves excellent performance in attribution tasks, demonstrating its superior ability to identify the intrinsic fingerprints of generative models. Our code is available at https://github.com/mwp-create-wonders/ExDA.

Xueyi Zhang 0001, Peiyin Zhu, Jinping Sui, Xiaoda Yang, Jiahe Tian, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Jun Tang 0001

The rapid evolution of deepfake techniques presents dual challenges for detection models: adapting to continuously shifting attack distributions while retaining previously learned knowledge. Although recent continual deepfake detection methods have made progress, they often rely on replay-based training, which limits scalability and deployment. Meanwhile, the task structure of deepfake detection offers a unique opportunity that remains under-explored: it is inherently a binary classification problem with a fixed label space, where the main difficulty lies in distributional drift rather than class expansion. This insight enables the modeling of each incremental distribution shift as a dedicated expert, focusing on specific forgery patterns. To this end, we propose a novel analytically driven, replay-free continual detection framework that eliminates the need for iterative gradient updates. In this framework, task-specific experts are constructed via closed-form ridge regression, requiring only a single forward pass and ensuring non-interference with previous tasks. To enhance the model's capacity for fine-grained forgery recognition, we introduce a lightweight Forgery-Aware Residual Enhancer (FARE). At inference, an Uncertainty-Guided Expert Selection module (UGES) dynamically routes each sample to the most confident expert, which does not require prior knowledge of the attack type. The proposed framework achieves a favorable trade-off between efficiency, privacy, and generalization. It achieves state-of-the-art performance across four benchmark datasets, with an average accuracy of 91.82% and only 1.78% forgetting. Notably, it improves cross-forgery generalization by 9.28% on unseen forgery types, demonstrating strong generalization.

Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan

The widespread availability of publicly accessible data on the internet accelerates the progress of deep learning but also raises concerns about unauthorized data usage for training neural networks. Early safeguard methods introduce small, carefully crafted perturbations via surrogate model into data to generate unlearnable data, aiming to prevent models from learning meaningful patterns. However, these methods lack robustness against adversarial training. Later, some works introduce adversarial examples to solve this problem but at the cost of increased overhead of the surrogate model. Recently, Convolution-based unlearnable data (CUDA), a surrogate-free method, has been proposed to address this issue by manually designed class-wise convolution kernels. Despite its success, CUDA suffers from high-frequency detail loss, perturbation hash collisions, and vulnerability to frequency filtering attacks. In this paper, we propose KBS (K-Space Bispectrum Steganography), which embeds class-specific information into the magnitude and phase components of the Fourier domain while preserving visual fidelity under reconstruction constraints. By directly performing steganography in the frequency domain, KBS preserves high-frequency details and avoids hash collisions with compact binary codes, enabling scalability to large-class datasets. Furthermore, KBS resists frequency filtering attacks by embedding perturbations in a way that remains imperceptible in the pixel space. Experimental results on public benchmarks demonstrate that KBS outperforms state-of-the-art methods.

Zhicheng Zhang 0002, Peizhuo Lv, Mengke Wan, Jiang Fang, Diandian Guo, Yezeng Chen, Yinlong Liu, Wei Ma, Jiyan Sun, Liru Geng

Recently, Deep Learning (DL) models have been increasingly deployed on end-user devices as On-Device AI, offering improved efficiency and privacy. However, this deployment trend poses more serious Intellectual Property (IP) risks, as models are distributed on numerous local devices, making them vulnerable to theft and redistribution. Most existing ownership protection solutions (e.g., backdoor-based watermarking) are designed for cloud-based AI-as-a-Service (AIaaS) and are not directly applicable to large-scale distribution scenarios, where each user-specific model instance must carry a unique watermark. These methods typically embed a fixed watermark, and modifying the embedded watermark requires retraining the model. To address these challenges, we propose Hot-Swap MarkBoard, an efficient watermarking method. It encodes user-specific n-bit binary signatures by independently embedding multiple watermarks into a multi-branch Low-Rank Adaptation (LoRA) module, enabling efficient watermark customization without retraining through branch swapping. A parameter obfuscation mechanism further entangles the watermark weights with those of the base model, preventing removal without degrading model performance. The method supports black-box verification and is compatible with various model architectures and DL tasks, including classification, image generation, and text generation. Extensive experiments across three types of tasks and six backbone models demonstrate our method's superior efficiency and adaptability compared to existing approaches, achieving 100% verification accuracy.

Eungi Lee, Jae Hyun Yoon, Seok Bong Yoo

Face-swapping deepfake poses significant risks, including privacy violations, misinformation, and defamation, amplified by the availability of pretrained models on open-source platforms. Proactive defense strategies aim to disrupt deepfake generation by modifying the original images to protect identity features. However, existing methods often introduce artifacts in facial images or rely on specific deepfake models, limiting their usability. To address these problems, we propose a style code orchestration in latent space (SCOL) method that obfuscates identity by fusing different identities in the latent space without requiring face recognition models. This study optimizes the generator to follow the original appearance while retaining the obfuscated identity via identity-preserving constraints. Further, appearance-dominant components in the latent code are aligned for visual consistency. An identity inversion attack is introduced using opposite style codes to improve the effectiveness of the defense. Experimental results demonstrate that SCOL robustly defends against various face-swapping deepfake methods, maintaining visual consistency.

Yitong Sun 0002, Yao Huang, Ruochen Zhang, Huanran Chen, Shouwei Ruan, Ranjie Duan, Xingxing Wei 0001

Despite the impressive generative capabilities of text-to-image (T2I) diffusion models, they remain vulnerable to generating inappropriate content, especially when confronted with implicit sexual prompts. Unlike explicit harmful prompts, these subtle cues, often disguised as seemingly benign terms, can unexpectedly trigger sexual content due to underlying model biases, raising significant ethical concerns. However, existing detection methods are primarily designed to identify explicit sexual content and therefore struggle to detect these implicit cues. Fine-tuning approaches, while effective to some extent, risk degrading the model's generative quality, creating an undesirable trade-off. To address this, we propose NDM, the first noise-driven detection and mitigation framework, which could detect and mitigate implicit malicious intention in T2I generation while preserving the model's original generative capabilities. Specifically, we introduce two key innovations: first, we leverage the separability of early-stage predicted noise to develop a noise-based detection method that could identify malicious content with high accuracy and efficiency; second, we propose a noise-enhanced adaptive negative guidance mechanism that could optimize the initial noise by suppressing the prominent region's attention, thereby enhancing the effectiveness of adaptive negative guidance for sexual mitigation. Experimentally, we validate NDM on both natural and adversarial datasets, demonstrating its superior performance over existing SOTA methods, including SLD, UCE, and RECE, etc.

Junlei Zhou, Jiashi Gao, Xinwei Guo, Haiyan Wu, Quanying Liu, Xiangyu Zhao 0001, Hongxin Wei, Xin Yao 0001, Xuetao Wei

Text-to-Image (T2I) diffusion models exhibit concerning tendencies to generate harmful imagery that perpetuates social biases and stereotypes, posing significant ethical risks in real-world applications. While existing mitigation approaches predominantly employ black-box methodologies through dataset augmentation or constrained fine-tuning, they face critical limitations, including high data acquisition costs and potential exacerbation of stereotypes during model retraining. Inspired by neuroscience principles where neurological dysfunction often stems from aberrant neural activation patterns, we propose a novel framework, StereoClinic, targeting the root cause of stereotype generation through direct neural intervention. Our solution introduces two synergistic components: Diffusion Deep Taylor Decomposition (DDTD) for precisely localizing stereotype-related neurons via Layer-wise Relevance Propagation (LRP) attribution analysis, and Stereotype Neuron Suppression (SNS) implementing targeted activation damping to neutralize bias propagation. Through extensive empirical evaluations across multiple bias dimensions, we demonstrate that our method achieves significant stereotype mitigation without compromising image quality or requiring additional training data. This neuro-inspired approach establishes a new paradigm for model interpretability and ethical alignment in generative AI systems.

Poyuan Mao, Cheng-Chang Tsai, Chun-Shien Lu

The great success of the diffusion model in image synthesis led to the release of gigantic commercial models, raising the issue of copyright protection and inappropriate content generation. Training-free diffusion watermarking provides a low-cost solution for these issues. However, the prior works remain vulnerable to rotation, scaling, and translation (RST) attacks. Although some methods employ meticulously designed patterns to mitigate this issue, they often reduce watermark capacity, which can result in identity (ID) collusion. To address these problems, we propose MaXsive, a training-free diffusion model generative watermarking technique that has high capacity and robustness. MaXsive best utilizes the initial noise to watermark the diffusion model. Moreover, instead of using a meticulously repetitive ring pattern, we propose injecting the X-shape template to recover the RST distortions. This design significantly increases robustness without losing any capacity, making ID collusion less likely to happen. The effectiveness of MaXsive has been verified on two well-known watermarking benchmarks under the scenarios of verification and identification.