论文检索

输入标题、作者或关键词,从 960 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
960篇论文匹配“Spectral Methods”
第 21 / 48 页

Jeonghyeok Do, Sungpyo Kim, Geunhyuk Youk, Jaehyup Lee, Munchurl Kim

PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment---caused by sensor placement, acquisition timing, and resolution disparity---induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11xfaster inference time and 0.63xthe memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.

Yang Li, Tingfa Xu, Shuyan Bai, Peifu Liu, Jianan Li 0001

Camouflaged Object Detection (COD) aims to identify objects that blend seamlessly into natural scenes. Although RGB-based methods have advanced, their performance remains limited under challenging conditions. Multispectral imagery, providing rich spectral information, offers a promising alternative for enhanced foreground-background discrimination. However, existing COD benchmark datasets are exclusively RGB-based, lacking essential support for multispectral approaches, which has impeded progress in this area. To address this gap, we introduce MCOD, the first challenging benchmark dataset specifically designed for multispectral camouflaged object detection. MCOD features three key advantages: (i) Comprehensive challenge attributes: It captures real-world difficulties such as small object sizes and extreme lighting conditions commonly encountered in COD tasks. (ii) Diverse real-world scenarios: The dataset spans a wide range of natural environments to better reflect practical applications. (iii) High-quality pixel-level annotations: Each image is manually annotated with precise object masks and corresponding challenge attribute labels. We benchmark eleven representative COD methods on MCOD, observing a consistent performance drop due to increased task difficulty. Notably, integrating multispectral modalities substantially alleviates this degradation, highlighting the value of spectral information in enhancing detection robustness. We anticipate MCOD will provide a strong foundation for future research in multispectral camouflaged object detection. The dataset is publicly accessible at https://github.com/yl2900260-bit/MCOD.

Yuntian Xiao, Shoulong Zhang, Zihang Zhang, Jiahao Cui 0001, Yan Wang, Shuai Li 0001

Generating highly realistic 4D interaction in real time is significant for visual content generation. Although existing works have validated to produce impressive dynamics by employing physical simulation and learned material mainly from pre-trained video diffusion models, it is still challenging to generate real-time 4D interaction with high-quality motion due to the heavy time consumption of the simulation solver and indirect material learning strategy. This paper proposes a novel physics-based 4D generation method, Phys4DRT, for arbitrary realistic real-time interaction on 3D Gaussian Splatting (3DGS) objects with direct motion supervision in time-frequency domain. Specifically, we devise a fast and differentiable eXtended Position Based Dynamics (XPBD) simulator as the light-weight controller for efficient physical evolution on a quasi-regular tetrahedral proxy mesh, into which we immerse the static 3DGS for efficient and stable deformation simulation. In addition, to learn the heterogeneous material for realistic motion, we directly supervise the generated dynamic 3D behavior by the motion representation of the optical flow and spectral volume extracted from the generated reference video, rather than indirect supervision in the color space used in previous approaches. We thoroughly conduct experiments on the public benchmarks to demonstrate the efficiency and effectiveness of our method. Our model can accelerate real-time 4D interaction generation by approximately x20 faster than the current Material Point Method (MPM) based approaches while achieving competitive visual quality compared with the state-of-the-art baselines.

Kunsheng Ma, Fan Qi, Changsheng Xu

The rapid development of music diffusion models has provided diverse paths for music creation transformations. However, existing methods still lack continuous strength regulation over stylistic attributes-specifically, they cannot achieve scalable adjustment of intensity (e.g., smooth transitions between ''gentle'' and ''intense'' jazz) while preserving spectral-temporal coherence. To address this, we propose RLScale-LoRA, a two-stage finetuning framework built on a structurally modified low-rank adaptation (LoRA) architecture with scale layers. In Stage 1, we finetune the modified LoRA to specialize in capturing attribute-aware latent spaces on unseen/seen music data. Stage 2 trains lightweight scale layers via proximal policy optimization (PPO), where reward functions enforce intermediate spectral-temporal state stability. Therefore, our RLScale-LoRA achieves precise, continuous music attribute transformations. Extensive experiments on Mtg-Jamendo and MedleyMD-Prompts datasets demonstrate RLScale-LoRA's superiority in granularity and coherence.

Shan Wang 0009, Weisi Lin, Yun Liu 0002, Libao Zhang

Unsupervised remote sensing dehazing remains a challenging and ill-posed task due to the absence of reliable supervision signals. Existing dehazing methods with unpaired data often oversimplify haze removal as style transfer, limiting generalization in complex scenarios. Moreover, current unimodal frameworks neglect cross-modal cues that could improve contextual reasoning. To address these issues, we propose a novel cross-modal guided self-supervised dehazing framework called CLIP-HNet, which achieves multi-model feature extraction, boundary-focused reconstruction and adaptive sample filtering. Specifically, to capture global-local contextual features, a hybrid feature interaction network is designed, which bridges the feature representations of multi models with global context-aware module (GCAM) and hybrid feature fusion module (HF2 M). Then, based on the hybrid features, a boundary-aware feature reconstruction (BFRec) is proposed to further refine edge details. Furthermore, a CLIP-guided progressive information distillation scheme is presented to dynamically prioritize training samples and distill useful signals, which predicts haze concentration by CLIP and progressively increases sample difficulty during the training stage. Finally, a frequency-domain texture matching (FTM) strategy refines texture and spectral details, enhancing the model's ability to recover fine details. Experiments on synthetic and real RSIs demonstrate that the proposed CLIP-HNet surpasses state-of-the-art approaches, achieving superior visual quality and quantitative performance.

Wanting Zhang, Jingxuan Zhang, Libao Zhang

Remote sensing image restoration under cloud and haze occlusions poses a significant challenge due to severe spectral degradation and spatial distortions. While recent generative models have shown promise in image restoration, they struggle with three key issues: (1) Lack of precise annotations, making supervised methods unreliable; (2) Unintended interference with clear regions, leading to distortion in unaffected areas; (3) Spectral and structural inconsistencies in heavily occluded regions, limiting realistic recovery. To address these challenges, we propose Saliency-Guided Adaptive Random Diffusion Strategy(SG-ARD), a novel blind restoration framework that integrates saliency-aware guidance with adaptive diffusion for enhanced reconstruction. First, we introduce a Saliency-Guided Pseudo-label Generation module (SGPG) to identify degraded regions and generate pseudo-labels for blind restoration. Second, we propose an Adaptive Random Diffusion Correction Strategy (ARDC), which employs a Random-Walk-based Diffusion and an Adaptive Enhancement module to refine local and global texture pseudo-labels. Lastly, we design a Spectral-Aware Consistency Loss (SAC) to improve spectral fidelity, ensuring that the generated content aligns with the real spectral distribution. Extensive experiments on three large-scale remote sensing datasets demonstrate that SG-ARD outperforms state-of-the-art generative restoration models, producing high-fidelity, visually coherent remote sensing images.

Teng Jin, Ziwen He, Zhangjie Fu, Songping Wang, Yueming Lyu, Yufei Shi

In recent years, adversarial attacks on video recognition models have attracted increasing attention. However, most existing strategies are extensions of image-based methods, where adversarial perturbations are computed independently and embedded into individual frames. This independent per-frame perturbation process wastes computational resources and leads to excessive query consumption. To address this problem, we introduce Frequency Domain Distributed Perturbations (FDP), a straightforward yet effective black-box video attack method using temporal correlations between video frames. Specifically, FDP first converts the input video into the frequency domain and calculates globally coordinated adversarial perturbations in the spectral space. By conducting global optimization in the frequency domain, FDP improves the effectiveness of each query, significantly decreasing the total number of queries needed. The resulting perturbations are temporally distributed across frames to preserve the spatiotemporal structure. Furthermore, we introduce a frequency-sensitive mask to identify the spectral regions most critical to the model's predictions. By applying perturbations only to these key frequency bands, FDP further reduces the perturbation search space and improves query efficiency. Extensive experiments demonstrate that our method significantly reduces query consumption while achieving higher attack success rates than state-of-the-art approaches.

Haitao Wang, Sijia Wen, Bo Guo

Monocular 3D Gaussian Splatting (3DGS) SLAM methods demonstrate outstanding performance in rapid dense 3D reconstruction. Yet former methods frequently exhibit suboptimal localization and mapping quality when processing indoor objects characterized by weak textures, dark colors, and high reflectivity (e.g., leather furniture), primarily due to insufficient surface feature information, even with the aid of depth sensors. To overcome these limitations, this work pioneers the integration of polarization information into the 3DGS SLAM framework. Specifically, we introduce a polarization integrated SLAM front-end that leverages the abundant planar features inherent in indoor environments. By incorporating a Chroma Boost mechanism, our approach effectively enhances the spectral multi-view consistency during the SLAM process, while the integration of a Gaussian-visible polarization difference improves the robustness of keyframe registration in low-texture scenarios. We further propose a flattened Gaussian regularization coupled with normal consistency constraints to capture the local geometric features of surfaces more accurately. Moreover, a novel integration of Pol-RGB hierarchical density plane segmentation and multi-scale plane self-constraint substantially enhances the quality of scene surface reconstruction, with further azimuth refinement achieved through the angle of linear polarization (AoLP). Extensive experiments demonstrate that, compared with previous SLAM methods, our approach significantly improves surface reconstruction quality.

Jinzhao Zhou, Zehong Cao, Yiqun Duan, Connor Barkley, Daniel Leong, Xiaowei Jiang, Quoc-Toan Nguyen, Ziyi Zhao, Thomas Do, Yu-Cheng Chang 等

This paper explores silent speech decoding in active brain-computer interface (BCI) systems, which offer more natural and flexible communication than traditional BCI applications. We collected a new silent speech dataset of over 120 hours of electroencephalogram (EEG) recordings from 12 subjects, capturing 24 commonly used English words for language model pretraining and decoding. Following the recent success of pretraining large models with self-supervised paradigms to enhance EEG classification performance, we propose Large Brain Language Model (LBLM) pretrained to decode silent speech for active BCI. To pretrain LBLM, we propose Future Spectro-Temporal Prediction (FSTP) pretraining paradigm to learn effective representations from unlabeled EEG data. Unlike existing EEG pretraining methods that mainly follow a masked-reconstruction paradigm, our proposed FSTP method employs autoregressive modeling in temporal and frequency domains to capture both temporal and spectral dependencies from EEG signals. After pretraining, we finetune our LBLM on downstream tasks, including word-level and semantic-level classification. Extensive experiments demonstrate significant performance gains of the LBLM over fully-supervised and pretrained baseline models. For instance, in the difficult cross-session setting, our model achieves 47.2% accuracy on semantic-level classification and 42.3% in word-level classification, outperforming baseline methods substantially. Our research advances silent speech decoding in active BCI systems, offering an innovative solution for EEG language model pretraining and a new dataset for fundamental research.

Siqi Song, Limin Yu, Jimin Xiao

Continual Learning (CL) enables models to sequentially acquire new knowledge while retaining previous knowledge. However, the challenge of catastrophic forgetting arises when new tasks interfere with previously acquired knowledge. Prompt-based approaches, leveraging pre-trained models, show promise in adapting to new tasks and reducing the risk of overfitting while mitigating catastrophic forgetting. However, existing approaches operate primarily in the spatial domain, neglecting the spectral entanglement between style-biased amplitude components and semantics-preserving phase components in feature representations. In this work, we propose the Spectral-Decomposed Prompting (SDP) method, a novel prompt-based approach that dynamically generates prompts based on the current input using a spectral decomposition strategy. By employing the Fast Fourier Transform (FFT), the query feature and the token embedding are transformed and decomposed into amplitude and phase spectra. SDP suppresses style-sensitive amplitude variations via spectral normalization while adaptively modulating phase components through task-aware attention mechanisms. It minimizes the disturbance of stylistic variations and enhances the semantic representations learning for prompts. Extensive experiments demonstrate that SDP significantly improves adaptability and performance in continual learning tasks, outperforming state-of-the-art methods while mitigating catastrophic forgetting.

Mengzu Liu, Junwei Xu, Tao Huang, Fangfang Wu, Le Dong, Xin Li 0005, Weisheng Dong

Multispectral image demosaicing aims to reconstruct full band multispectral images from a compressed spectral mosaic images. Although existing learning-based methods have made progress in multispectral image demosaicing, there still exist intrinsic performance bottlenecks due to the heavy undersampling according to mosaic pattern. To address this issue, we propose Polarity memory network with quant attention to establish global correlation, thus reconstructing high-quality multispectral images from compressed spectral mosaic images. Our proposed Polarity memory network adaptively encapsulates reconstruction-oriented representations, then amplifies relevant ones and reducing noise from irrelevant ones in a polarity-aware manner to better cater to the enhancement of different spectral information with linear computational complexity. Moreover, considering existing methods' inability to adequately compensate for long distance interactions in reconstruction, we introduce a quant attention paradigm that categorize tokens into semantic-aware groups using an efficient quant operation for attention computation. Experimental results show our method achieves state-of-the-art performance on various simulation datasets and better vision results on real-world datasets.

Guyue Jin, Tianming Zhao 0003, Jiacan Yan, Tian Tian 0006

Multi-modal feature fusion under conditions of image misalignment remains a significant challenge in multispectral object detection. Existing approaches predominantly rely on cross-attention mechanisms; however, when local features are sparse, inadequate feature capture hinders accurate alignment and results in distorted fusion outcomes. To address this problem, we propose a novel multispectral fusion detection network, CSSFDet, which leverages the intrinsic correlations among image regions in visual recognition to dynamically enhance local features via global semantic constraints during the fusion process. Specifically, we introduce a Contextual Region Feature Fusion Module (CRFM) that regulates the fusion process through a selective state-space formulation, adaptively incorporating surrounding context to compensate for local feature degradation caused by misalignment. Moreover, we design a Complementary Enhancement Module (CoE) to ensure both distinctiveness and completeness of modality-specific features. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across multiple datasets, attaining 84.1% mAP50 on the DroneVehicle dataset-a 20% improvement over the baseline. It also shows strong performance on the misaligned CVC-14 dataset, and sensitivity analysis on data shifts further underscores its robustness to misalignment.

Xuyao Liu, Jiahui Qu, Wenqian Dong

Currently, fusion-based hyperspectral image super-resolution (fusion-based HSI-SR) has become an efficient technology to improve the spatial resolution of hyperspectral images. However, in real scenarios, it may not be possible to obtain high-resolution multispectral images (HR-MSI) of the same temporal and region corresponding to low-resolution hyperspectral images (LR-HSI) due to the limitations of imaging conditions and environmental changes. In view of this spatial-temporal constraint, it becomes a feasible solution to regard HR-MSI, which has similar spatial structure and semantics to LR-HSI, as a reference to assist in reconstruction. Therefore, this paper proposes a Cross-Correlation & Self-Similarity Guided Texture Transfer Network (C2S2TNet), which utilizes the texture details of HR-MSI and the self-similarity information of LR-HSI to achieve reference-based hyperspectral image super-resolution. Specifically, we design a Cross-Correlation & Self-Similarity Guided Cluster-Aware Matching (C2S2CAM) strategy, which realizes multi-correspondence texture matching and feature aggregation in non-local regions based on dynamic clustering and cluster-aware graph structure, effectively alleviating the misuse and underuse of information. In addition, we also propose a Spectral-Spatial State-Space Fusion Module (S2-SSFM) based on the state-space model to perform feature fusion and enhancement in both spatial and spectral domains to ensure that the target HR-HSI maintains the spatial-spectral structural consistency with the LR-HSI. Experimental verification shows that C2S2TNet can achieve excellent performance in cross-temporal and cross-regional scenarios, confirming the effectiveness of this method. Code can be accessed at https://github.com/Jiahuiqu/C2S2TNet.

Lamei Di, Bin Zhang 0022, Yiming Wang, Wenxia Zhang

Salient object detection in optical remote sensing images (ORSI-SOD) faces unique challenges due to complex backgrounds, diverse scales, and multi-directional objects. Existing methods primarily rely on visual features, often struggling to distinguish salient objects from visually similar backgrounds. To address this limitation, we leverage large language models (LLMs) to expend existing ORSI-SOD datasets with detailed textual annotations, creating a more comprehensive benchmark for image-text ORSI-SOD. Building upon this foundation, we propose the Frequency Meets Semantics Network (FMS-Net), a novel framework that integrates text-visual fusion with directional spectral enhancement for ORSI-SOD. FMS-Net consists of two key innovations: the Hierarchical Multi-Modal Dual-Channel Fusion (HMDF) module and the Adaptive Directional Spectral Enhancement (ADSE) module. The HMDF module enables bidirectional interactions between visual and textual features via parallel global-local attention mechanisms, progressively enriching visual representations with semantic context. Meanwhile, the ADSE module enhances feature representations in the frequency domain, capturing directional patterns and boundary details critical for accurate saliency detection. Extensive experiments on two public datasets, ORSSD and EORSSD, demonstrate that FMS-Net outperforms state-of-the-art methods, particularly in complex scenes with ambiguous boundaries. Our work paves the way for integrating multi-modal and frequency-based approaches in the interpretation of optical remote sensing images (ORSI).

Xueheng Li, Xuanhua He, Tao Hu 0027, Jie Zhang 0033, Man Zhou 0003, Chengjun Xie, Yingying Wang 0005, Bo Huang 0001

Pan-sharpening aims to improve the spatial resolution of low-resolution multispectral (LRMS) image by integrating high-frequency information from corresponding texture-rich panchromatic (PAN) image. While RWKV architecture has demonstrated remarkable global perception with linear computational efficiency in vision tasks, its inherent sequential scanning mechanism critically compromises local spatial coherence, hindering high-frequency reconstruction. To bridge this gap, we tailor Freq-RWKV, the first spatial-frequency adaptive RWKV featuring dual-domain scanning where wavelet-guided path selection dynamically modulates scanning granularity and orientation according to spectral-spatial information density distributions. Building upon this innovation, we architect the hierarchical U-shaped fusion network that strategically coordinates granularity-aware scanning across spatial and frequency domains, enabling adaptive trade-offs between performance and complexity. The U-shaped architecture implements coarse-to-fine enhancement: in the encoding stage, Coarse Structural Interaction (CSI-RWKV) module preserves geometric dependencies via window-constrained recurrent scanning while encoding structural priors into LRMS features; during decoding, the Fine-grained Frequency Interaction (FFI-RWKV) module performs edge-aware refinement through differentiable frequency-adaptive window partitioning, where multi-scale spectral wavelet attention prioritizes high-frequency PAN components extracted via discrete wavelet transform (DWT). This hybrid decomposition strategy maintains spectral integrity through approximation coefficients while detail coefficients regulate frequency-gated fusion thresholds. Extensive experiments on multiple satellite datasets validate the effectiveness of the proposed method.

Qiyin Zhong, Xianglin Qiu, Xiaolei Wang, Zhen Zhang, Gang Liu, Jimin Xiao

Multimodal Anomaly Detection (MMAD) has attracted significant attention in industrial defect inspection as it can simultaneously leverage the complementary information from different modalities to achieve higher-precision detection. Among existing MMAD approaches, dual-branch reverse distillation is widely adopted because of its efficiency in avoiding large-scale data storage. However, it suffers from two key issues. First, the alignment of cross-modal features can lead to a loss of modality-specific characteristics. Second, when one modality indicates normal while another shows anomalies, anomaly detection may be misled by that modality ambiguity. To address these challenges, we propose a Frequency-Aware Multimodal Reverse Distillation (FAMRD) framework from the frequency domain perspective. Specifically, we introduce a frequency spectral feature alignment module that aligns the low- and medium-frequency components across modalities to preserve global shape consistency, while maintaining high-frequency modality-specific details. In addition, we design a frequency spectral anomaly synthesis module. It perturbs the normal feature of one modality to create modality consistent anomalies, fuses it with another modality normal feature to mimic modality ambiguous anomalies, and adds them to the reverse distillation process for decision boundary optimization. Extensive experiments on standard MMAD benchmarks demonstrate that FAMRD achieves competitive performance in both anomaly detection and localization, outperforming state-of-the-art methods.

Mufan Liu, Wu Ran, Zhiquan He, Zuojie Xie, Hong Lu 0001, Peirong Ma

Low-light image enhancement aims to improve brightness, suppress noise, and recover accurate color and structure, requiring precise illumination modeling and reliable reflectance recovery. However, most Retinex-based methods adopt explicit, multi-stage pipelines prone to decomposition bias, error accumulation, and chromatic entanglement between illumination and reflectance. To tackle these issues, we propose IDAR (Implicit Decomposition, illumination Adjustment, and reflectance Restoration), a unified Retinex-inspired framework with two key innovations. First, we design an implicit decomposition strategy based on dual-branch feature learning: a low-frequency-constrained illumination branch models lighting with chromaticity awareness, while a contrast-guided reflection branch preserves details by decoupling reflectance from illumination. This implicit design avoids intermediate supervision and reduces decomposition bias. Second, we introduce the Illumination Chromaticity Expansion Module (ICEM), which employs text-guided chromaticity learning to enhance chromaticity perception. By learning a reflectance-independent spectral representation, ICEM reduces color shifts and improves fidelity under complex lighting. Experiments on multiple benchmarks validate the superior visual quality, quantitative performance, and physical interpretability of IDAR.

Junwei Zhu, Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001

Spatial-spectral fusion offers a promising alternative to expensive equipment in high-resolution hyperspectral (HrHs) imaging. However, training separate models for different scaling factors remains costly. To address this, we propose the Arbitrary-scale Fusion Neural Operator (AFNO), a lightweight solution for HrHs fusion across arbitrary scalings. Instead of entities, AFNO treats low-resolution hyperspectral (LrHs) and high-resolution multispectral (HrMs) images as functions and performs meticulously designed integrations as the mapping operator. The key components include Attention-Driven Convolution Integration (ADCI) to restore discretization invariance disrupted by convolutions, Implicit Neural Functional Integration (INFI) for cross-domain interaction of spatial degradations, and Galerkin-type Integration as a decoder for high-frequency details. Additionally, the bonded activation opeartor are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, AFNO holds significantly better generalization on arbitrary scaling factors, yet requiring only 0.07M parameters.

Wei Li 0034, Junwei Zhu, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001

By clustering pixels with locally similar values, superpixel-based approaches have shown great potential in processing hyperspectral images (HSI) , thereby reducing the computational burden associated with large spatial dimensions. However, specific for spatial-spectral fusion (SSF), superpixel segmentation is inherently non-differentiable and irreversible; hence it is inapplicable. To address the issues, we propose a semantic transformer-based solver, namely SpecSolver, which is basically inspired by the benefits of superpixel-based approaches, yet with the inner mechanism completely improved. The core idea lies in learning the intrinsic semantic states of HSIs hidden behind discretized pixel representations. Specifically, we propose a new Semantic-Attention to adaptively split the image domain into a series of learnable slices of flexible shapes, where image pixels under similar semantic states will be ascribed to the same slice. By calculating attention to the Semantic-Superpixel tokens encoded from slices, SpecSolver can effectively capture intricate semantic correlations from the vast number of pixels, which also empowers the solver with an endogenous capacity for modeling different magnification scales and allows for efficient computation in linear complexity. On that basis, we elaborate a SpatialNet module, which extracts multiscale local spectral information, and a FreqNet module, which supplements global information, capturing subtle details and variations across different spectra. Experiments on two benchmark SSF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies validate the mentioned contributions.

Peirong Zhang 0001, Kai Ding 0009, Lianwen Jin

In this paper, we propose SPECTRUM, a temporal-frequency synergistic model that unlocks the untapped potential of multi-domain representation learning for online handwriting verification (OHV). SPECTRUM comprises three core components: (1) a multi-scale interactor that finely combines temporal and frequency features through dual-modal sequence interaction and multi-scale aggregation, (2) a self-gated fusion module that dynamically integrates global temporal and frequency features via self-driven balancing. These two components work synergistically to achieve micro-to-macro spectral-temporal integration. (3) A multi-domain distance-based verifier then utilizes both temporal and frequency representations to improve discrimination between genuine and forged handwriting, surpassing conventional temporal-only approaches. Extensive experiments demonstrate SPECTRUM's superior performance over existing OHV methods, underscoring the effectiveness of temporal-frequency multi-domain learning. Furthermore, we reveal that incorporating multiple handwritten biometrics fundamentally enhances the discriminative power of handwriting representations and facilitates verification. These findings not only validate the efficacy of multi-domain learning in OHV but also pave the way for future research in multi-domain approaches across both feature and biometric domains. Code is publicly available at https://github.com/NiceRingNode/SPECTRUM.