论文检索

输入标题、作者或关键词,从 1,767 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,767篇论文匹配“Energy”
第 35 / 89 页

Xiubo Liang, Hongzhi Wang 0009, Zigen Li, Jinxing Han, Yu Zhao, Weidong Geng

Spiking Neural Network (SNN), as a next-generation neural network technology, use binary spike signals as carriers of information. They offer advantages such as low energy consumption, low computational complexity, and high information transmission rates. However, deep SNNs suffer from the gradient vanishing problem due to issues like the non-differentiability of the step function and neuron dormancy. To address this gradient problem, we first propose the MCLIF neuron, which optimizes the backpropagation mechanism and compensates for gradient information from both temporal and spatial dimensions. Furthermore, we design a spiking attention mechanism tailored to the temporal characteristics of SNNs. By introducing the QK Memory to embed temporal features, we make full use of information from different time steps. Additionally, we propose a gradient correction module to enhance the model's representational power from both temporal and spatial dimensions. The proposed SGM-Transformer achieves state-of-the-art (SOTA) performance in image classification tasks such as CIFAR10, CIFAR100, and CIFAR10-DVS, and also excels in industrial defect classification scenarios. The code will be made available after the paper is accepted.

Qian Sun 0014, Chengzhuo Lu, Wenyu Chen 0001, Wenjie Wei, Jingya Wang, Jieyuan Zhang, Xiaoli Liu, Yalan Ye, Yang Yang 0002, Malu Zhang

Spiking Neural Networks (SNNs) have garnered significant attention due to their biological plausibility and low power consumption. While spiking transformers enhance performance by combining SNNs with transformer architecture, most rely on rate coding, limiting energy efficiency. Temporal coding methods, such as Time-To-First-Spike (TTFS) coding, offer a more efficient alternative by encoding information based on the timing of a single spike. However, integrating TTFS with transformer architecture faces challenges due to incompatibility with batch normalization (BN) and residual connections (RC), which disrupt the precise spike firing times. In this paper, we propose temporal-coded BN (tBN) and temporal-coded RC (tRC) to address these issues. Building on tBN and tRC, we develop temporal-coded spiking attention (TSA) and temporal-coded spiking transformer (T-SpikeFormer), the first to combine TTFS coding with transformer architecture. Experimental results show our model achieves state-of-the-art performance for temporal-coded SNNs and comparable results to rate-coded SNNs while significantly reducing power consumption.

Yongzheng Liu, Siru Zhong, Gefeng Luo, Weilin Ruan, Yuxuan Liang 0002

The rapid urbanization process has significantly increased building energy consumption and carbon emissions, making reliable electricity load forecasting crucial for energy management. However, accurate load forecasting faces three key challenges: (1) complex impact of multimodal data, (2) inter-building semantical relationships, and (3) uncertainty modeling of load patterns. To address these, we propose MMLoad, a novel diffusion-based multimodal framework for multi-scenario building load forecasting with three innovations: (i) a Multimodal Data Enhancement Pipeline generating rich building descriptions using LLMs and integrating temporal factors to analyze multimodal impacts; (ii) a Cross-modal Relation Encoder discovering latent interdependencies through hierarchical fusion, projecting buildings into a unified spatio-temporal (ST) embedding space; and (iii) a Scenario-Conditioned Diffusion Generator employing transformer-based denoising with Scenario-Adaptive Normalization (SAN) for diverse trajectory generation with uncertainty quantification. Experiments show MMLoad outperforms state-of-the-art baselines in accuracy while generating plausible future scenarios, establishing a new paradigm for multimodal learning in smart energy systems.

Ziyu Wang, Yiming Du, Rui Ning, Lusi Li

Incomplete multi-view clustering (IMVC) deals with real-world scenarios where certain views are partially missing, posing significant challenges to effective clustering. Most existing IMVC approaches face a trade-off: imputation-free methods suffer from information bias and imbalance, while full-imputation methods risk introducing and propagating noise. To overcome these limitations, we propose Energy-Based Deep Incomplete Multi-View Clustering (Energy-DIMC), a novel selective-imputation framework that leverages energy-based models (EBMs) to guide reliable imputations and robust clustering. EBMs assess data compatibility by assigning lower energy to more coherent structures, effectively modeling complex inter-view and inter-sample dependencies. Inspired by EBMs, Energy-DIMC integrates four key components: 1) a view feature projector that learns view-specific features and projects them into a common feature space; 2) an energy-guided selective imputation module that identifies the most reliable source view for each view based on view energies, and performs feature imputation only when cross-view transfer is feasible, avoiding unreliable imputations; 3) an energy-based representation fusion module that aggregates observed and selectively imputed features across views via a view attention mechanism, generating view-coherent representations; 4) an energy-enhanced contrastive alignment module that enforces consistency between view-specific and view-coherent representations using dual-level energy signals to preserve true positives. Extensive experiments demonstrate that Energy-DIMC outperforms state-of-the-art IMVC methods across diverse missing-view scenarios. The code is available at https://github.com/sunway677/EnergyIMVC.

Naichuan Zheng, Yuchen Du, Hailun Xia, Zeyu Liang 0001

For skeleton-based action recognition, Graph Convolutional Networks (GCNs) are effective models. Still, their reliance on floatingpoint computations leads to high energy consumption, limiting their applicability in battery-powered devices. While energy-efficient, Spiking Neural Networks (SNNs) struggle to model skeleton dynamics, leading to suboptimal solutions. We propose Signal-SGN (Spiking Graph Convolutional Network), which utilizes the temporal dimension of skeleton sequences as the spike time steps and represents features as multi-dimensional discrete stochastic signals for temporal-frequency domain feature extraction. It combines the 1D Spiking Graph Convolution (1D-SGC) module and the Frequency Spiking Convolution (FSC) module to extract features from the skeleton represented as spiking form. Additionally, the Multi-Scale Wavelet Transform Feature Fusion (MWTF) module is proposed to extract dynamic spiking features and capture frequency-specific characteristics, enhancing classification performance. Experiments across three large-scale datasets reveal Signal-SGN exceeding state-of-the-art SNN-based methods in accuracy and computational efficiency while attaining comparable performance with GCN methods and significantly reducing theoretical energy consumption.

Xiangping Zheng, Xuan Feng, Bo Wu 0026, Bin Ren, Wei Li 0109, Xiuxin Hao, Xun Liang 0001, Bin Tang, Zhiwen Yu 0001

Cross-domain graph anomaly detection (GAD) aims to identify nodes that significantly deviate from normal patterns in unseen target domains, showing great potential in applications such as multimedia content security and financial risk control. However, existing methods often rely on semantic information trained on individual datasets, which makes it difficult to capture node commonalities across domains and limits generalization in complex multimedia environments. To address these challenges, we propose Zero-GAD, a universal Zero-shot Graph Anomaly Detection framework tailored for cross-domain scenarios. Zero-GAD leverages a novel de-semanticized strategy to train a unified detection model that can be directly applied to unseen domains without retraining or fine-tuning. The framework is built upon two key components: (1) a Global Information Unification Module, which projects graph data into the spectral domain and performs normalization to align the energy distribution in the frequency space; and (2) a Node-Neutralized Discrepancy Scoring Module that leverages the discrepancy between the original and reconstructed node representations to produce effective anomaly scores. Extensive experiments show that Zero-GAD achieves superior accuracy compared to existing models under a GAD setting.

Yingbing Liu, Fei Ma 0001, Yanan Wu, Xinxin Zuo, Fan Zhang 0007, Yang Wang 0003

Generalized category discovery (GCD) aims to group unlabeled samples from known and unknown classes when only part of the labeled data in the known classes is given. It allows the model to adapt to dynamic environments by discovering novel categories. However, when we applied the GCD approach to the decentralized open world, we still encountered the following challenges: (1) none of labeled data easily obtained in the open world, (2) heterogeneous label spaces across different environments, (3)representation degradation caused by fine-tuning models with limited data in specific environments. To address the above challenges, we introduce a new and practical task, namely Cloud-edge GCD (CE-GCD). Different from semi-supervised GCD, CE-GCD assumes that we only have a base model trained on common public categories, and aims to perform personalized unsupervised novel category discovery in multiple environments with heterogeneous label spaces. Data from different environments or clients cannot be shared, only model parameters can be transferred. To tackle this problem, we propose a novel GCD framework based on energy-guided known class discrimination and multi-level contrastive learning. In each client, we first use the classifier of the base model to distinguish between known and unknown classes, and then perform unsupervised learning on the unknown classes. Each client transfers category information through prototypes to assist learning. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach.

Wei Wu 0011, Shiqi Li, Ling Chen 0006, Fangfang Li 0004, Chuan Luo 0002

Real-world networks, particularly those in web and social media, are dynamic with evolving node attributes and structures, often involving billions of nodes and edges. Dynamic attributed network embedding is a powerful tool for capturing these changes, enabling data owners and problem owners to better understand interactions and trends for more effective engagement and decision-making. While some existing algorithms are capable of handling very large-scale dynamic attributed networks with billions of nodes and edges, they often suffer from accuracy loss or high computational overhead. In this paper, we propose a practical and sustainable framework of sketching very large-scale dynamic attributed networks called VLS2ketch, which incorporates incremental embedding updates alongside storage-efficient, binarized representation of both node attributes and topological variations. By the sparse random projection technique in an incremental update manner, VLS2ketch significantly reduces the energy-intensive computational workload while maintaining accuracy. Also, we introduce an information decay mechanism, which adapts to temporally varying topologies and node attributes. This mechanism ensures that outdated information gradually diminishes over time. Extensive experiments on real-world very large-scale datasets demonstrate that our proposed VLS2ketch method delivers comparable embedding quality against the state-of-the-art learning-based competitors with dramatically reduced runtime. We have released the source code and the datasets in https://github.com/AIandBD/graph-hashing/tree/main/VLS2ketch. .

Kun Xie 0010, Renchi Yang, Sibo Wang 0001

Clustering over a graph seeks to partition the nodes therein into disjoint groups such that nodes within the same cluster are tightly-knit, while those across clusters are distant from each other. In practice, graphs are often attended with rich attributes, which are termed attributed graphs. By leveraging the complementary nature of graph topology and node attributes in such graphs, graph neural networks (GNNs) have obtained encouraging performance in graph clustering. However, existing GNN-based approaches strongly rely on the homophilic assumption of the input graph, and thus, largely fail on heterophilic graphs and others embodying numerous missing or noisy links, which are widely present in real life. To bridge this gap, this paper presents DGAC, an effective graph-agnostic solution for graph clustering. Particularly, DGAC overcomes the limitations of prior works by exploiting the high-order connectivity of nodes within not only the input graph G but also the affinity graph H underlying the attribute data. To achieve this goal, we first unify the embedding and clustering generations into a coherent framework optimizing the Dirichlet Energy on both G and H. Based thereon, theoretically-grounded solvers are developed for efficient constructions of the embeddings and clusters via graph diffusion operations, which aggregate features from specific neighbors, enabling the capture of high-order semantics from G or H. On top of that, DGAC includes three training loss functions that facilitate effective feature extraction and clustering. Extensive experiments, comparing DGAC against 12 baselines over 12 homophilic or heterophilic graph datasets, showcase that DGAC consistently and considerably outperforms all competitors in terms of clustering quality measured against ground truth labels.

Niloy Talukder, Croix Gyurek, Mohammad Al Hasan

With the adoption of deep learning models to low-power, small-memory edge devices, energy consumption and storage usage of such models have become a key concern. The problem exacerbates even further with ever-growing data and equally-matched bulkier models. This concern is particularly pronounced for graph data due to its quadratic storage, irregular (non-grid) geometry, and very large size. Typical graph data, such as road networks, infrastructure networks, and social networks, easily exceeds millions of nodes, and several gigabytes of storage is needed just to store the node embedding vectors, let alone the model parameters. In recent years, the memory issue has been addressed by moving away from memory-intensive double precision floating-point arithmetic towards single-precision or even half-precision, often by trading-off marginally small performance. Along this effort, we propose Node2Binary, which embeds graph nodes in as few as 128 binary bits, thereby reducing the memory footprint of vertex embedding vectors by several orders of magnitude. Node2Binary. leverages a fast community detection algorithm to convert the given graph into a hierarchical partition tree and then find embeddings of graph vertices in binary space by solving a combinatorial optimization (CO) task over the tree edges. CO is NP-hard, but Node2Binary uses an innovative combination of discrete gradient descent and randomization to solve this task effectively and efficiently. Extensive experiments over four real-world graphs show that Node2Binary achieves competitive performance compared to the state-of-the art graph embedding methods in both node classification and link prediction tasks.

Yu Feng 0015, Yangli-ao Geng, Yifan Zhu 0001, Zongfu Han, Xie Yu, Kaiwen Xue 0001, Haoran Luo 0001, Mengyang Sun, Guangwei Zhang 0003, Meina Song

Federated learning (FL) has gained widespread attention for its privacy-preserving and collaborative learning capabilities. Due to significant statistical heterogeneity, traditional FL struggles to generalize a shared model across diverse data domains. Personalized federated learning addresses this issue by dividing the model into a globally shared part and a locally private part, with the local model correcting representation biases introduced by the global model. Nevertheless, locally converged parameters more accurately capture domain-specific knowledge, and current methods overlook the potential benefits of these parameters. To address these limitations, we propose PM-MoE architecture. This architecture integrates a mixture of personalized modules and an energy-based personalized modules denoising, enabling each client to select beneficial personalized parameters from other clients. We applied the PM-MoE architecture to nine recent model-split-based personalized federated learning algorithms, achieving performance improvements with minimal additional training. Extensive experiments on six widely adopted datasets and two heterogeneity settings validate the effectiveness of our approach. The source code is available at https://github.com/dannis97500/PM-MOE.

Yuxuan Liang 0002, Yu Zheng 0004, Chuishi Meng, Yanhua Li, Jieping Ye, Philip S. Yu, Ouri Wolfson

The swift advancement of urbanization has resulted in the growth of numerous large cities, which have enhanced the lives of many individuals but have also created significant challenges, such as air pollution, higher energy consumption, and traffic congestion. Addressing these issues was nearly unfeasible in the past due to the intricate and ever-changing nature of urban environments. Today, however, advancements in sensing technologies and extensive computing infrastructures have generated vast amounts of big data related to urban areas, including information on human mobility, air quality, traffic patterns, and geographic data. Inspired by the potential for creating smarter cities, we developed a vision for urban computing that seeks to harness insights from diverse and extensive data collected in urban settings, using this valuable information to tackle the critical problems our cities currently encounter.

Di Yang, Yanhai Xiong

Path planning is a critical challenge for Autonomous Underwater Vehicles (AUVs) due to complex underwater environments, including ocean currents, dynamic obstacles, and limited sensing capabilities. The lack of a standardized benchmarking framework has hindered direct comparisons between algorithms, slowing progress in the field. To address this, we introduce an open-source benchmarking platform for underwater AUV path planning, designed to provide a unified evaluation environment, automated performance assessment, and reproducible experiments. Built on the HoloOcean simulation platform, our benchmark incorporates realistic underwater dynamics, such as ocean currents, static and dynamic obstacles, and sensor models. It supports a range of path planning tasks, from basic obstacle avoidance to complex scenarios with current disturbances. The platform is compatible with classical algorithms (e.g., A*, RRT), evolutionary methods (e.g., GA, ACO), and deep reinforcement learning (e.g., Soft Actor-Critic, SAC). We define key evaluation metrics, including path efficiency (length, smoothness, energy consumption), task success rate, collision rate, and computational cost. Automated tools enable systematic algorithm comparisons across scenarios, generating standardized performance results and visualizations. This open-source, extensible framework aims to advance underwater path planning research by enabling fair comparisons and guiding future algorithm development. It provides a scalable foundation for evaluating AUV path planning methods under simulated real-world conditions, fostering innovation in the field. All source code and experimental configurations will be available on the GitHub Page: https://github.com/IoET-y/UP-bench.

Zhe Li 0011, Xiangfei Qiu, Peng Chen 0038, Yihang Wang 0004, Hanyin Cheng, Yang Shu 0001, Jilin Hu, Chenjuan Guo, Aoying Zhou, Christian S. Jensen 等

Time Series Forecasting (TSF) is key functionality in numerous fields, such as financial investment, weather services, and energy management. Although increasingly capable TSF methods occur, many of them require domain-specific data collection and model training and do not generalize well when applied in other domains. Time Series Foundation Models (TSFMs) that are pre-trained on massive heterogeneous time series data aim to overcome these limitations. The prospects for generalizability have spurred the development of a new generation of TSFMs. This study proposes a benchmark, TSFM-Bench, to facilitate comprehensive and unified evaluation of TSFMs. TSFM-Bench covers a wide range of TSFMs, including those based on large language models and those pre-trained on time series data. TSFM-Bench supports multiple forecasting scenarios, including zero-shot, few-shot, and full-shot, enabling assessment across the full range of adaptation strategies. TSFM-Bench also provides a standardized experimental protocols for critical evaluation processes such as dataset splitting, loading, normalization, and few-shot sampling, facilitating consistency and fairness. We report on an extensive evaluation of TSFMs across a diverse range of datasets spanning multiple domains and exhibiting varied statistical characteristics. Specifically, we identify pros and cons and inherent limitations of existing TSFMs, and we propose potential directions for new model designs.

Shiyu Wang 0001, Wei Lu 0030, Jiawei Li 0017, Xiaoming Shi 0001, Xinyue Zhong, Zhou Ye 0001, Ming Jin 0005, Qingsong Wen

Many real-world applications contain data in the form of multivariate time series (TS) with the hierarchical structure, where classic methods forecasting each TS independently are inadequate for coherency (i.e., satisfying the hierarchical aggregation constraints). Furthermore, the discrepancies between statistical properties of different levels can be huge, exacerbated by non-Gaussian distributions and non-linear correlations. In this paper, we propose a novel end-to-end hierarchical TS forecasting model, i.e., a Flow-based Reconcile Transformer (FRT). FRT employs a conditional normalizing flow-based autoregressive transformer, to represent complex data distribution, while simultaneously reconciling the forecasts to ensure coherency. Go beyond other state-of-the-art methods, FRT accomplishes forecasting and reconciliation simultaneously, while avoiding any post-processing steps. Moreover, FRT is a deep model that does not rely on any strong assumptions such as unbiased estimates or Gaussian distribution. Our experiments are conducted on four real-world hierarchical datasets from different industrial domains (three public ones and a dataset from the application servers of our company's data center) and the results demonstrate the efficacy of our proposed method. Our method has been implemented extensively within the production environments of a prominent global payment company. It has emerged as a cornerstone for workload forecasting within their data center and plays a critical role in the optimization of cloud computing resource allocation across the entire cluster. This successful deployment of the application demonstrates that our approach achieves precise hierarchical workload prediction, which is of great significance for efficient resource scheduling, enhancing resource utilization, and reducing server resource use and energy consumption in data centers.

Keyue Shi, Qianqian Shen, Zhongda Qi, Junyao Yang, Zhaoming Ye, Jiajun Bu, Haishuai Wang

Dual-energy X-ray absorptiometry (DXA) enables accurate bone mineral density but requires specialized equipment and protocols. X-ray-based BMD screening offers opportunistic early detection, though prior methods struggle with X-ray intensity variations and demand large datasets. We introduce a cross-modal knowledge distillation BMD prediction framework (CMKD-BMD) fusing X-ray/CT data to enhance X-ray-only BMD prediction. Each single-modal student network employs a multi-scale visual extractor for hierarchical features, an unsupervised graph-based structural learner for anatomical relationships, and an adaptive fusion module to generate a unified representation. The teacher network integrates single-modal representations from students and transfers multimodal knowledge from the teacher to students. Our model outperforms existing methods on both the collected dataset (comprising 1620 X-rays and 280 CT cases) and the publicly available VerSe2019 dataset, demonstrating superior BMD estimation performance. Furthermore, we developed OrthoSim, an orthopedic surgical simulation platform with CMKD-BMD, which has shown promising clinical effectiveness in trial evaluations. Our code is available at https://github.com/KeyueShi/CMKD-BMD.

Adrien Petralia, Philippe Charpentier, Youssef Kadhi, Themis Palpanas

Millions of smart meters have been deployed worldwide, collecting the total power consumed by individual households. Based on these data, electricity suppliers offer their clients energy monitoring solutions to provide feedback on the consumption of their individual appliances. Historically, such estimates have relied on statistical methods that use coarse-grained total monthly consumption and static customer data, such as appliance ownership. Non-Intrusive Load Monitoring (NILM) is the problem of disaggregating a household's collected total power consumption to retrieve the consumed power for individual appliances. Current state-of-the-art (SotA) solutions for NILM are based on deep-learning (DL) and operate on subsequences of an entire household consumption reading. However, the non-stationary nature of real-world smart meter data leads to a drift in the data distribution within each segmented window, which significantly affects model performance. This paper introduces NILMFormer, a Transformer-based architecture that incorporates a new subsequence stationarization/de-stationarization scheme to mitigate the distribution drift and that uses a novel positional encoding that relies only on the subsequence's timestamp information. Experiments with 4 real-world datasets show that NILMFormer significantly outperforms the SotA approaches. Our solution has been deployed as the backbone algorithm for EDF's (Electricité De France) consumption monitoring service, delivering detailed insights to millions of customers about their individual appliances' power consumption.

Tengfei Lyu, Jindong Han, Hao Liu 0026

Nuclear radiation, which refers to the energy emitted from atomic nuclei during decay, poses significant risks to human health and environmental safety. Recently, advancements in monitoring technology have facilitated the effective recording of nuclear radiation levels and related factors, such as weather conditions. The abundance of monitoring data enables the development of accurate and reliable nuclear radiation forecasting models, which play a crucial role in informing decision-making for individuals and governments. However, this task is challenging due to the imbalanced distribution of monitoring stations over a wide spatial range and the non-stationary radiation variation patterns. In this study, we introduce NRFormer, a novel framework tailored for the nationwide prediction of nuclear radiation variations. By integrating a non-stationary temporal attention module, an imbalance-aware spatial attention module, and a radiation propagation prompting module, NRFormer collectively captures complex spatio-temporal dynamics of nuclear radiation. Extensive experiments on two real-world datasets demonstrate the superiority of our proposed framework against 11 baselines. NRFormer has been deployed online to provide 1-24-day nuclear radiation forecasts, empowering individuals and governments with timely, data-driven decisions for emergency response and public safety. Our framework is designed for general applicability and can be readily adapted for deployment in other regions. The deployed system is available at https://NRFormer.github.io and the dataset and code of the predictive model are available at https://github.com/usail-hkust/NRFormer.

Haoye Chai, Shiyuan Zhang, Xiaoqian Qi, Baohua Qiu, Yong Li 0008

Mobile traffic forecasting allows operators to anticipate network dynamics and performance in advance, offering substantial potential for enhancing service quality and improving user experience. It involves multiple tasks, including long-term prediction, short-term prediction, and generation tasks that do not rely on historical data. By leveraging the different types of mobile network data generated from these tasks, operators can perform a variety of network optimizations and planning activities, such as base station (BS) deployment, resource allocation, energy optimization, etc. However, existing models are often designed for specific tasks and trained with specialized data, and there is a lack of universal models for traffic forecasting across different urban environments. In this paper, we propose a Universal model for Mobile traffic forecasting (UoMo), aiming to handle diverse forecasting tasks of short/long-term predictions and distribution generation across multiple cities to support network planning and optimization. UoMo combines diffusion models and transformers, where various spatio-temporal masks are proposed to enable UoMo to learn intrinsic features of different tasks, and a contrastive learning strategy is developed to capture the correlations between mobile traffic and urban contexts, thereby improving its transfer learning capability. Extensive evaluations on 9 real-world datasets demonstrate that UoMo outperforms current models in various forecasting tasks and zero/few-shot learning. It shows an average accuracy improvement of 27.85%, 18.57%, and 15.6% in long-term prediction, short-term prediction, and generation tasks, respectively, showcasing its strong forecasting capability. We deploy UoMo on China Mobile's JiuTian platform, leveraging the predicted mobile data to optimize live networks. This optimization includes BS deployment, resulting in a 25.3% increase in served users, and BS sleep control, which reduces equipment depreciation by 40.7%. The source code is available online: https://github.com/tsinghua-fib-lab/UoMo.

Lin Zuo, Yongqi Ding, Mengmeng Jing, Kunshan Yang, Biao Chen, Yunqian Yu

This paper explores the application of spiking neural networks (SNNs), known for their low-power binary spikes, to bearing fault diagnosis, bridging the gap between high-performance AI algorithms and real-world industrial scenarios. In particular, we identify two key limitations of existing SNN fault diagnosis methods: inadequate encoding capacity that necessitates cumbersome data preprocessing, and non-spike-oriented architectures that constrain the performance of SNNs. To alleviate these problems, we propose a Multi-scale Residual Attention SNN (MRA-SNN) to simultaneously improve the efficiency, performance, and robustness of SNN methods. By incorporating a lightweight attention mechanism, we have designed a multi-scale attention encoding module to extract multiscale fault features from vibration signals and encode them as spatio-temporal spikes, eliminating the need for complicated preprocessing. Then, the spike residual attention block extracts high-dimensional fault features and enhances the expressiveness of sparse spikes with the attention mechanism for end-to-end diagnosis. In addition, the performance and robustness of MRA-SNN is further enhanced by introducing the lightweight attention mechanism within the spiking neurons to simulate the biological dendritic filtering effect. Extensive experiments on MFPT, JNU, Bearing, and Gearbox benchmark datasets demonstrate that MRA-SNN significantly outperforms existing methods in terms of accuracy, energy consumption, and noise robustness, and is more feasible for deployment in real-world industrial scenarios. Our codes are available at https://github.com/yqding326/MRA-SNN.