论文检索

输入标题、作者或关键词,从 12,319 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
12,319篇论文匹配“Datasets and Benchmarks”
第 196 / 616 页

Yunyu Zou, Yishu Liu 0001, Jun Liang 0002, Bingzhi Chen

Cross-Domain Few-Shot Learning (CD-FSL) aims to transfer knowledge acquired from a source domain with abundant data to the target domain with limited labeled samples. Recent advancements have enhanced model generalization through Perturbation Augmentation (PA), facilitating more effective knowledge transfer. However, PA-based CD-FSL methods still suffer from two critical challenges, i.e., (1) limited diversity of augmented samples, making it difficult to cover the true distribution of unseen domains, and (2) conflicting gradients during model optimization, where augmented and original samples drive the model's optimization in opposing directions. To address these issues, we propose a novel PA-based framework with Style-Decoupled Augmentation (SDA) and Gradient-Conflict Adjustment (GCA) for Cross-Domain Few-Shot Learning, which is termed ''SG-FSL''. Specifically, SDA decouples the source domain style into style weights and basis styles, generating diverse unseen styles by perturbing the style weights to reweight the basis styles. Meanwhile, GCA leverages the angular relationships between the domain-specific gradient directions of augmented and original features, adaptively adjusting the gradient directions of original features to ensure that the model acquires diverse domain knowledge without interference, guiding it toward conflict-free optimization. Comprehensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our method over state-of-the-art baselines.

Yu-Wei Zhan, Fan Liu 0008, Xin Luo 0006, Xin-Shun Xu, Liqiang Nie, Mohan Kankanhalli

Human-Object Interaction (HOI) detection involves detecting human-object pairs and predicting their interactions. However, it faces significant challenges due to the complexity of human behavior and the diverse contexts in which interactions occur. Contextual cues, such as the participants involved, body language, and the surrounding environment, are crucial for accurately identifying interactions, particularly those that are ambiguous or previously unseen. In this paper, we propose ConCue, a novel approach that integrates contextual cue generation with feature extraction to enhance HOI detection. Specifically, we design specialized prompts tailored for Large Vision-Language Models (VLMs), enabling the generation of rich contextual cues from images. These cues are then seamlessly integrated into HOI detection through a feature extraction module with a multi-tower architecture we developed, which effectively incorporates contextual information into both instance and interaction detection processes. Extensive experimental results demonstrate the effectiveness of ConCue. Integrating ConCue with state-of-the-art HOI methods leads to significant performance improvements on two widely used benchmark datasets, highlighting the potential of our approach in advancing HOI detection.

Zhicong Wu, Hongbin Xu, Gang Xu, Ping Nie, Zhixin Yan, Jinkai Zheng, Liangqiong Qu, Ming Li 0073, Liqiang Nie

Recent advancements in Generalizable Gaussian Splatting have enabled robust 3D reconstruction from sparse input views by utilizing feed-forward Gaussian Splatting models, achieving superior cross-scene generalization. However, while many methods focus on geometric consistency, they often neglect the potential of text-driven guidance to enhance semantic understanding, which is crucial for accurately reconstructing fine-grained details in complex scenes. To address this limitation, we propose TextSplat-the first text-driven Generalizable Gaussian Splatting framework. Specifically, our framework employs three parallel modules to obtain complementary representations: the Diffusion Prior Depth Estimator for accurate depth information, the Semantic Aware Segmentation Network for detailed semantic information, and the Multi-View Interaction Network for refined cross-view features. Then, in the Text-Guided Semantic Fusion Module, these representations are integrated via the text-guided and attention-based feature aggregation mechanism, resulting in enhanced 3D Gaussian parameters enriched with detailed semantic cues. Experimental results on various benchmark datasets demonstrate improved performance compared to existing methods across multiple evaluation metrics, validating the effectiveness of our framework. The code will be publicly available.

Shun Zou, Yi Zou, Juncheng Li 0003, Guangwei Gao, Guo-Jun Qi

Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular rain patterns and complex geometric overlaps challenge single-paradigm architectures, necessitating a unified framework to integrate complementary global-local and spatial-channel representations. To address this, we propose a novel Cross Paradigm Representation and Alignment Transformer (CPRAformer). Its core idea is the hierarchical representation and alignment, leveraging the strengths of both paradigms (spatial-channel and global-local) to aid image reconstruction. It bridges the gap within and between paradigms, aligning and coordinating them to enable deep interaction and fusion of features. Specifically, we use two types of self-attention in the Transformer blocks: sparse prompt channel self-attention (SPC-SA) and spatial pixel refinement self-attention (SPR-SA). SPC-SA enhances global channel dependencies through dynamic sparsity, while SPR-SA focuses on spatial rain distribution and fine-grained texture recovery. To address the feature misalignment and knowledge differences between them, we introduce the Adaptive Alignment Frequency Module (AAFM), which aligns and interacts with features in a two-stage progressive manner, enabling adaptive guidance and complementarity. This reduces the information gap within and between paradigms. Through this unified cross-paradigm dynamic interaction framework, we achieve the extraction of the most valuable interactive fusion information from the two paradigms. Extensive experiments demonstrate that our model achieves state-of-the-art performance on eight benchmark datasets and further validates CPRAformer's robustness in other image restoration tasks and downstream applications.

Zhiqian Xia, Haifeng Xia, Shichao Jin, Wei Wang 0335, Zhengming Ding, Xiaochun Cao

Point cloud completion is crucial for downstream tasks in 3D visual perception. However, existing methods often struggle to generalize to real-world scans due to their heavy reliance on abundant paired point clouds for training and their neglect of the distribution shift between training and testing datasets. To address these limitations, this paper explores a practical and challenging setting: ''source-free domain adaptive point cloud completion'', where a well-trained source model must adapt to the target data distribution without access to source data, aiming to improve completion performance. To tackle this problem, we propose a novel method called ''Dual-Stage Preservation and Fusion'' (DSPF), which comprises two key training stages tailored to this new setting. In the source preservation stage, we introduce graph structural alignment and marginal feature alignment to preserve and transfer essential knowledge from the source domain. In the target fusion stage, we design a self-supervised loss to capture the geometric structure of target instances and establish a bidirectional interaction mechanism to transfer partial source knowledge to the target distribution. Extensive experiments on various cross-domain point cloud completion benchmarks demonstrate that our proposed DSPF significantly outperforms existing methods, validating its effectiveness and robustness in source-free domain adaptation scenarios. Our code is available at https://github.com/ZhiXia-SEU/DSPF.

Tung-I Chen, Dae Yeol Lee, Guan-Ming Su, Mohammad Hajiesmaili, Ramesh K. Sitaraman

We present NIVM, a lightweight and efficient view morphing framework that learns coordinate transforms between views, enabling real-time, user-controlled perspective shifts on resource-constrained devices. Unlike existing view interpolation methods that compromise visual quality or require high data overhead, NIVM integrates seamlessly into multi-view video streams as compact metadata per frame, enabling the synthesis of high-quality intermediate views and interactive transitions from sparse viewpoints. To avoid dependence on explicit 3D geometry, which may be unavailable, we introduce a dual-branch training strategy: a teacher network operates in rectified stereo space to supervise the morpher in the original image domain. By inheriting the monotonicity constraints of epipolar geometry, our morphing network produces visually plausible pixel flows while avoiding the reprojection artifacts prevalent in depth-based methods. Compared to recent pose-free sparse-view Gaussian Splatting approaches, NIVM achieves competitive results without the need to construct or transmit volumetric representations. Experiments show that NIVM achieves the lowest memory footprint, highest inference efficiency, and top-tier visual quality across multiple benchmark datasets.

Yishu Liu 0001, Zhiming Chen, Desen Wang, Xiaoling Luo 0001, Bingzhi Chen, Guangming Lu 0002

Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel concepts from limited training samples without forgetting previously encountered classes. Recent advancements have leveraged Parameter-Efficient Tuning (PET) strategies on pre-trained models to enhance FSCIL performance. However, current PET-based FSCIL approaches still suffer from the challenges posed by catastrophic collapse of general prompt and limited adaptability of specific prompt . To this end, we redefine the function of the PET paradigm with both gradient-aware prompting (GAP) and router-free adapters (RFA) to boost the performance of FSCIL, termed as "PET-GPRA". To dynamically balance the retention of previously learned general knowledge and the acquisition of novel class information across sessions, the GAP paradigm adaptively adjusts the updated gradient of the general prompt by leveraging the angular relationship between the general knowledge gradient and the novel knowledge gradient. Meanwhile, the RFA mechanism utilizes the semantic similarity between class attributes to replace the routing network, guiding the integration of adapter information, in which adapters serve as specific prompts to enhance the adaptability. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our proposed PET-GPRA framework over state-of-the-art baselines.

Xueyi Zhang 0001, Jialu Sun, Chengwei Zhang, Xianghu Yue, Tianfang Xiao, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001

Event cameras, with their microsecond-level temporal resolution and sparse visual encoding, provide a transformative paradigm for automatic lip reading (ALR). However, event data inherently lack explicit spatial structure and exhibit a pronounced frequency-domain bias. The low-frequency components fail to capture crucial lip structural information, which fundamentally impedes the modeling of intra-frame topological dependencies and inter-frame semantic evolution-both of which are critical for robust lip reading. To this end, we propose FAST-HG, a Frequency-Aware SpatioTemporal HyperGraph framework specifically designed for event-based lip reading. First, we apply low-frequency perturbation to improve the model's robustness for capturing discriminative features, and integrate adaptive high-frequency filtering to enhance edge-aware representations. Then, we construct a Spatial Region Hypergraph (SRH) and a Temporal Semantic Hypergraph (TSH). The former captures intra-frame topological dependencies among lip regions, while the latter explicitly models inter-frame structural associations throughout the lip movement process, enabling the model to capture discriminative patterns in lip dynamics. Furthermore, we propose a viseme-aware label smoothing strategy, where a novel viseme-level edit distance is designed to quantify visual similarities between classes and guide the construction of soft labels. FAST-HG achieves 79.85% and 84.03% accuracy on the DVS-Lip and DVS-LRW100 datasets, respectively, significantly outperforming prior methods and establishing a new benchmark for event-based lip reading.

Xihang Hu, Fuming Sun, Jiazhe Liu, Feilong Xu, Xiaoli Zhang 0001

Semi-supervised Camouflaged Object Detection (SSCOD) aims to reduce reliance on costly pixel-level annotations by leveraging limited annotated data and abundant unlabeled data. However, existing SSCOD methods based on Teacher-Student frameworks suffer from severe prediction bias and error propagation under scarce supervision, while their multi-network architectures incur high computational overhead and limited scalability. To overcome these limitations, we propose ST-SAM, a highly annotation-efficient yet concise framework that breaks away from conventional SSCOD constraints. Specifically, ST-SAM employs Self-Training strategy that dynamically filters and expands high-confidence pseudo-labels to enhance a single-model architecture, thereby fundamentally circumventing inter-model prediction bias. Furthermore, by transforming pseudo-labels into hybrid prompts containing domain-specific knowledge, ST-SAM effectively harnesses the Segment Anything Model's potential for specialized tasks to mitigate error accumulation in self-training. Experiments on COD benchmark datasets demonstrate that ST-SAM achieves state-of-the-art performance with only 1% labeled data, outperforming existing SSCOD methods and even matching fully supervised methods. Remarkably, ST-SAM requires training only a single network, without relying on specific models or loss functions. This work establishes a new paradigm for annotation-efficient SSCOD. Codes will be available at https://github.com/hu-xh/ST-SAM.

Hanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao, Tianyu Fu 0001, Minghui Wu, Beiping Tan, Huiying Li 0002

Visual selective attention, driven by individual preferences, regulates human prioritization of visual stimuli by bridging subjective cognitive mechanisms with objective visual elements, thereby steering the semantic interpretation and hierarchical processing of dynamic visual scenes. However, existing models and datasets predominantly neglect the influence of subjective cognitive diversity on fixation behavior. Conventional saliency prediction models, typically employing segmentation approaches, rely on low-resolution imagery to generate saliency heatmaps, subsequently upscaled to native resolutions, which limiting their capacity to capture personalized attention patterns. Furthermore, MLLMs are constrained by factors such as hallucinations, making it very costly to strictly adhere to the expected format in tasks involving multiple point predictions, and achieving precise point positioning is challenging. To address these limitations, we present Subjective Personalized Attention for Ad vertisement Videos, namely SPA-ADV, a large-scale multimodal dataset capturing gaze behaviors from over 4,500 participants varying in age and gender with 486 videos. Furthermore, we propose PRE-MAP, a novel eye-tracking saliency model that characterizes Personalized visual disparities through Reinforcement learning-optimized Eye-tracking, built upon MLLMs and guided by Multi-Attribute user profiles to predict Points. To ensure MLLMs produce prediction points that are both format-correct and spatially accurate, we introduce Consistency Group Relative Policy Optimization (C-GRPO), inspired by the variability in eye movement points and Multi-Attribute profiles. Extensive experiments on SPA-ADV and other benchmarks demonstrate the effectiveness of our approach. The code and dataset are available at https://github.com/mininglamp-MLLM/PRE-MAP.

Huynh Dang Nguyen, Trong-Thang Pham, Ngan Le, Van Nguyen

The electrocardiogram (ECG) is an essential and effective tool for diagnosing heart diseases. However, its effectiveness can be compromised by noise or unavailability of one or more leads of the standard 12-lead recordings, resulting in diagnostic errors or uncertainty. To address these challenges, we propose TolerantECG, a foundation model for ECG signals that is robust to noise and capable of functioning with arbitrary subsets of the standard 12-lead ECG. TolerantECG training combines contrastive and self-supervised learning frameworks to jointly learn ECG signal representations alongside their corresponding knowledge-retrieval-based text report descriptions and corrupted or lead-missing signals. Comprehensive benchmarking results demonstrate that TolerantECG consistently ranks as the best or second-best performer across various ECG signal conditions and class levels in the PTB-XL dataset, and achieves the highest performance on the MIT-BIH Arrhythmia Database. The source is available at this link: https://github.com/Fsoft-AIC/TolerantECG

Fansheng Zeng, Bineng Zhong 0001, Haiying Xia, Yufei Tan, Xiantao Hu, Liangtao Shi, Shuxiang Song 0001

Contextual reasoning with constraints is crucial for enhancing temporal consistency in cross-frame modeling for visual tracking. However, mainstream tracking algorithms typically associate context by merely stacking historical information without explicitly supervising the association process, making it difficult to effectively model the target's evolving dynamics. To alleviate this problem, we propose RSTrack, which explicitly models and supervises context reasoning via three core mechanisms. 1) Context Reasoning Mechanism : Constructs a target state reasoning pipeline, converting unconstrained contextual associations into a temporal reasoning process that predicts the current representation based on historical target states, thereby enhancing temporal consistency. 2) Forward Supervision Strategy : Utilizes true target features as anchors to constrain the reasoning pipeline, guiding the predicted output toward the true target distribution and suppressing drift in the context reasoning process. 3) Efficient State Modeling : Employs a compression-reconstruction mechanism to extract the core features of the target, removing redundant information across frames and preventing ineffective contextual associations. These three mechanisms collaborate to effectively alleviate the issue of contextual association divergence in traditional temporal modeling. Experimental results show that RSTrack achieves state-of-the-art performance on multiple benchmark datasets while maintaining real-time running speeds. Our code is available at https://github.com/GXNU-ZhongLab/RSTrack.

Lizhi Xiong, Peipeng Yu, Yue Wu

Perceptual hashing has garnered significant attention for its wide-ranging applications in image retrieval and authentication domains. However, existing algorithms often struggle to detect subtle manipulations confined to small regions of an image. In this paper, we introduce a novel framework, Manipulation-Aware Deep Perceptual Hashing (MADPHash), which leverages feature consistency to enhance sensitivity to such subtle manipulations. MADPHash explicitly treats tampered images as a distinct category, incorporates a tampering detection objective into the perceptual hash generation process, and employs a Consistency Constraint Module to amplify discrepancies between tampered and untampered regions. Comprehensive experiments conducted on five benchmark datasets demonstrate that MADPHash significantly improves the detection of subtle manipulations while maintaining robustness against content-preserving transformations, outperforming several state-of-the-art perceptual hashing methods.

Yu Tong, Weihai Lu, Xiaoxi Cui, Yifan Mao, Zhejun Zhao

Lately, the academic community has been showing growing interest in multi-domain fake news detection, and in particular, incorporating multimodal information into this field has emerged as a highly promising research direction. However, existing methods often struggle with: (1) Insufficient intrinsic domain adaptation during representation Learning; (2) Amplified negative transfer from entangled domain style and content representations; and (3) Neglecting domain-varying modality uncertainty. To address these issues, we propose Domain-Aware Prompt Tuning (DAPT), an innovative framework for multimodal multi-domain fake news detection. DAPT leverages Multimodal Prompt Tuning for parameter-efficient domain adaptation of pretrain models. An Adaptive Domain Debias Module disentangles domain features from veracity signals guided by content to mitigate negative transfer. Furthermore, inspired by the Variational Information Bottleneck, an Uncertainty-Aware Multimodal Fusion mechanism adaptively aggregates modalities based on domain-specific reliability. Extensive experiments demonstrate that DAPT significantly outperforms state-of-the-art baselines on benchmark datasets.

Chengcheng Xing, Yanyu Xu 0001, Yonghui Xu, Lizhen Cui 0001

Unified Anomaly Detection (UAD) aims to identify anomalies across diverse domains without access to target domain data during training. Unlike traditional anomaly detection methods that rely on training separate models for each domain, UAD employs a single model to generalize across multiple categories. A key challenge lies in the domain shift between seen and unseen data, which requires capturing invariant discriminative patterns between reference and query images across different domains during in-context learning for unified anomaly detection. To tackle this, we propose a novel UAD framework to learn the invariant discriminative patterns through pre-, in- and post-processing modules. First, a pre-processing VLM-guided data augmentation module generates diverse and semantically consist images, followed by a latent-space filtering mechanism. Second, an in-processing Adaptive VQ memory module stores representative discriminative patterns to enable robust residual comparison. Third, a post-processing GUR (Geometric distributions Upgrade Representation) feature augmentation module models geometric feature distributions to synthesize informative prompts, improving the quality of feature delta estimation for anomaly scoring. Extensive experiments on benchmark datasets demonstrate that our method achieves superior generalization in detecting anomalies across unseen domains, outperforming existing state-of-the-art approaches.

Xiang Huang 0004, Ao Luo, Xiao Wu 0001, Zhaoquan Yuan

Human-Object Interaction (HOI) detection serves a broad spectrum of applications. Despite significant progress, current approaches encounter difficulties in effectively handling Non-Contact Human-Object Interaction (NCHOI) scenarios, where humans and objects remain physically apart. To address these challenges, this paper proposes a novel approach, named Latent Interactiveness Field Modeling (LIFM), which enhances HOI detection by capturing long-range contextual dependencies. Specifically, the Latent Interactiveness Field (LIF) is introduced to define potential interactive relationships between humans and objects. To complement this, the LIF Fusion Encoder is designed to adaptively fuse visual features with LIF, resulting in more informative and discriminative feature representations. The Mobile Scanning HOI Dataset (MSHD) is introduced as a comprehensive benchmark to systematically assess the robustness of existing methods on both common HOI and NCHOI in real-world applications. Extensive experimentation indicates that the proposed approach outperforms existing state-of-the-art techniques. It offers substantial improvements, particularly in NCHOI scenarios, which highlight its effectiveness in resolving issues related to long-range interactions.

Jinhao Li 0001, Zijian Chen 0001, Runze Jiang, Tingzhu Chen, Changbo Wang, Guangtao Zhai

The oracle bone inscription (OBI) recognition plays a significant role in understanding the history and culture of ancient China. However, the existing OBI datasets suffer from a long-tail distribution problem, leading to biased performance of OBI recognition models across majority and minority classes. With recent advancements in generative models, OBI synthesis-based data augmentation has become a promising avenue to expand the sample size of minority classes. Unfortunately, current OBI datasets lack large-scale structure-aligned image pairs for generative model training. To address these problems, we first present the Oracle-P15K, a structure-aligned OBI dataset for OBI generation and denoising, consisting of 14,542 images infused with domain knowledge from OBI experts. Second, we propose a diffusion model-based pseudo OBI generator, called OBIDiff, to achieve realistic and controllable OBI generation. Given a clean glyph image and a target rubbing-style image, it can effectively transfer the noise style of the original rubbing to the glyph image. Extensive experiments on OBI downstream tasks and user preference studies show the effectiveness of the proposed Oracle-P15K dataset and demonstrate that OBIDiff can accurately preserve inherent glyph structures while transferring authentic rubbing styles effectively. The dataset, code, and pre-trained models are available at https://github.com/LJHolyGround/Oracle-P15K.

Quanhong Peng, Dan Zhang 0016, Dong Zhao, Jianpeng Zhang, Meihua Song, Chenlei Lv

For camera-based image capturing, the impact of exposure or camera parameters (ISO sensitivity, shutter speed, and aperture F-number) on imaging quality is decisive. Such parameters interact in a coupled manner during the imaging process to determine the exposure quality and the degree of blur in a photograph. Naturally, decoupling such parameters from images holds significant value for applications like image quality assessment and illumination optimization. However, there has been no systematic research dedicated to this topic. In this paper, we propose a new benchmark, Cam-Bench, for estimating camera parameters on images directly. It collects an image dataset Cam-10K with various indoor scenes and accurate labels of camera parameters. Based on Cam-10K, we propose a camera parameter estimation network to decouple and regress recorded exposure information. To the best of our knowledge, Cam-Bench is the first benchmark for camera parameter estimation. Experiments demonstrate that it can enhance the performance of various downstream applications.The source code has been made publicly available at: https://github.com/pengquanhong/CamBench.

Yanrui Yu, Tianfei Zhou, Jiaxin Sun, Lianpeng Qiao, Lizhong Ding 0003, Ye Yuan 0001, Guoren Wang

In modern urban environments, camera networks generate massive amounts of operational footage -- reaching petabytes each day -- making scalable video analytics essential for efficient processing. Many existing approaches adopt an SQL-based paradigm for querying such large-scale video databases; however, this constrains queries to rigid patterns with predefined semantic categories, significantly limiting analytical flexibility. In this work, we explore a language-driven video analytics paradigm aimed at enabling flexible and efficient querying of high-volume video data driven by natural language. Particularly, we build Lava, a system that accepts natural language queries and retrieves traffic targets across multiple levels of granularity and arbitrary categories. Lava comprises three main components: 1) a multi-armed bandit-based efficient sampling method for video segment-level localization; 2) a video-specific open-world detection module for object-level retrieval; and 3) a long-term object trajectory extraction scheme for temporal object association, yielding complete trajectories for object-of-interests. To support comprehensive evaluation, we further develop a novel benchmark by providing diverse, semantically rich natural language predicates and fine-grained annotations for multiple videos. Experiments on this benchmark demonstrate that Lava improves F1-scores for selection queries by 14% reduces MPAE for aggregation queries by 0.39, and achieves top-k precision of 86% while processing videos 9.6x faster than the most accurate baseline. Our code and dataset are available at https://github.com/yuyanrui/LAVA.

Zhaohu Xing, Lihao Liu, Tian Ye 0001, Sixiang Chen, Yijun Yang, Guang Liu 0006, Lei Zhu 0003

Current video mirror detection models demonstrate satisfactory performance by analyzing different attributes of mirrors and incorporating temporal information. However, these models still struggle to detect mirrors in complex and dynamic scenarios. A simple yet critical visual cue is that objects reflected in a mirror appear to be farther away than the mirror itself. Motivated by this observation, some studies propose to explicitly analyze the Depth of Mirror (DOM) to effectively localize mirrors - DOM refers to distinct perceived distances that make mirror regions appear farther away from their surroundings. However, merely analyzing the DOM is insufficient in some scenes where the object behind the mirror also appears distant. Meanwhile, the changes in the DOM across different video frames are also important for video mirror detection, yet this aspect has not been fully explored. To address these issues, we devise a novel framework called FTM-Net, which includes two main contributions: a Pattern-Compensated DOM estimation strategy and a Dual-Granularity Affinity module. The Pattern-Compensated DOM estimation strategy uses multiple visual mirror patterns to refine the DOM, enhancing the accuracy of mirror localization in a single image. Furthermore, the Dual-Granularity Affinity module can effectively detect mirrors in video sequences by tracking and integrating DOM changes across video frames. Experimental results on two benchmark datasets show that our model significantly outperforms other state-of-the-art methods in the video mirror detection task. We shall release our trained models, code, and results.