论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 20 / 81 页

Yaojun Wu 0001, Chaoyi Lin, Yiming Wang 0008, Semih Esenlik, Zhaobin Zhang, Kai Zhang 0007, Li Zhang 0006

This paper explores the application of enhancement filtering techniques in neural video compression. Specifically, we categorize these techniques into in-loop contextual filtering and out-of-loop reconstruction enhancement based on whether the enhanced representation affects the subsequent coding loop. In-loop contextual filtering refines the temporal context by mitigating error propagation during frame-by-frame encoding. However, its influence on both the current and subsequent frames poses challenges in adaptively applying filtering throughout the sequence. To address this, we introduce an adaptive coding decision strategy that dynamically determines filtering application during encoding. Additionally, out-of-loop reconstruction enhancement is employed to refine the quality of reconstructed frames, providing a simple yet effective improvement in coding efficiency. To the best of our knowledge, this work presents the first systematic study of enhancement filtering in the context of conditional-based neural video compression. Extensive experiments demonstrate a 7.71% reduction in bit rate compared to state-of-the-art neural video codecs, validating the effectiveness of the proposed approach.

Yuheng Wu 0006, Thanh-Tung Nguyen, Lucas Liebe, Quang Tau, Pablo Espinosa Campos, Jinghan Cheng, Dongman Lee

With the rapid proliferation of the Internet of Things, video analytics has become a cornerstone application in wireless multimedia sensor networks. To support such applications under bandwidth constraints, learning-based adaptive quantization for video compression has demonstrated strong potential in reducing bitrate while maintaining analytical accuracy. However, existing frameworks often fail to fully exploit the fine-grained quality control enabled by modern blockbased video codecs, leaving significant compression efficiency untapped. In this paper, we present How2Compress, a simple yet effective framework designed to enhance video compression efficiency through precise, fine-grained quality control at the macroblock level. How2Compress is a plug-and-play module and can be seamlessly integrated into any existing edge video analytics pipelines. We implement How2Compress on the H.264 codec and evaluate its performance across diverse real-world scenarios. Experimental results show that How2Compress achieves up to 50.4% bitrate savings and outperforms baselines by up to 3.01× without compromising accuracy, demonstrating its practical effectiveness and efficiency.Code is available at https://github.com/wyhallenwu/how2compress and a reproducible docker image at https://hub.docker.com/r/wuyuheng/how2compress

Fenghao Tian, Mingtao Feng, Jianqiao Luo, Zijie Wu, Longlong Mei, Lijie Yang, Weisheng Dong, Yaonan Wang 0001

Fine-grained cross-view localization seeks to predict ground-level camera positions within GPS-tagged aerial images by matching ground and aerial views. Existing methods often rely on large-scale ground truth annotations from specific regions, but performance degrades due to domain shifts when models trained in one area are applied to another. However, collecting region-specific annotations for each area is costly or infeasible. To address this, we propose a self-distillation curriculum learning framework that generalizes pretrained localization models to unseen new areas. Our approach introduces a Dirichlet-based quality assessment strategy to evaluate teacher-generated pseudo labels, where high uncertainty signals noisy predictions and low uncertainty indicates clean samples. This uncertainty is used to guide an easy-to-hard curriculum learning strategy, where easy samples are prioritized initially, and more challenging samples are progressively incorporated, enabling effective student training. Furthermore, we develop a joint optimization scheme that updates both the student model and pseudo labels, applying adaptive label smoothing to mitigate label noises and taking full advantage of new area data. Extensive experimental results on the VIGOR and KITTI benchmarks demonstrate that our method outperforms state-of-the-art approaches in new area localization, achieving superior accuracy without additional supervision.

Chunyu Qiao, Tong Liu, Yucheng Zhang, Zhiwei Fan, Pengjin Xie, Zhen Wang 0071, Liang Liu 0001

In large-scale short-video platforms, CDN resource selection plays a critical role in maintaining users' Quality of Experience (QoE) while controlling escalating traffic costs. To better understand this phenomenon, we conduct in-the-wild network measurements during video playback in a production short-video system. The results reveal that CDNs delivering higher average QoE often come at greater financial cost, yet their connection quality fluctuates even within a single video-underscoring a fundamental and dynamic trade-off between QoE and cost. However, the problem of sustaining high QoE under cost constraints remains insufficiently investigated in the context of CDN selection for short-video streaming. To address this, we propose PIRA, a dynamic resource selection algorithm that optimizes QoE and cost in real-time during video playback. PIRA formally integrating QoE and cost by a mathematical model, and introduce a intra-video control-theoretic CDN resource selection approach which can balance QoE and cost under network dynamics. To reduce the computation overheads, PIRA employs state-space pruning and adaptive parameter adjustment to efficiently solve the high-dimensional optimization problem. In large-scale production experiments involving 450,000 users over two weeks, PIRA outperforms the production baseline, achieving a 2.1% reduction in start-up delay, 15.2% shorter rebuffering time, and 10% lower average unit traffic cost, demonstrating its effectiveness in balancing user experience and financial cost at scale.

Sarmistha Das 0001, R. E. Zera Marveen Lyngkhoi, Sriparna Saha 0001, Alka Maurya

The dynamic propagation of social media has broadened the reach of financial advisory content through podcast videos, yet extracting insights from lengthy, multimodal segments (30-40 minutes) remains challenging. We introduce FASTER(Financial Advisory Summariser with Textual Embedded Relevant images), a modular framework that tackles three key challenges: (1) extracting modality-specific features, (2) producing optimized, concise summaries, and (3) aligning visual keyframes with associated textual points. FASTER employs BLIP-2 for semantic visual descriptions, OCR for textual patterns, and Whisper-based transcription with Speaker diarization as BOS features. A modified Direct Preference Optimization (DPO)-based loss function, equipped with BOS-specific fact-checking, ensures precision, relevance, and factual consistency against the human-aligned summary. A ranker-based retrieval mechanism further aligns keyframes with summarized content, enhancing interpretability and cross-modal coherence. To acknowledge data resource scarcity, we introduce Fin-APT, a dataset comprising 470 publicly accessible financial advisory pep-talk videos for robust multimodal research. Comprehensive cross-domain experiments confirm FASTER's strong performance, robustness, and generalizability when compared to Large Language Models (LLMs) and Vision-Language Models (VLMs). By establishing a new standard for multimodal summarization, FASTER makes financial advisory content more accessible and actionable, thereby opening new avenues for research.

Haizhou Wang, Guobing Zou, Fei Xu 0009, Yangguang Cui, Tongquan Wei

Federated learning (FL), an emerging data-secure distributed training paradigm, unites massive isolated Internet of Things (IoT) device nodes to collaboratively train a global neural network (NN) model without the exposure of their local multimedia data. However, constrained by the synchronous NN model integration nature of FL, there is a training latency inconsistency among heterogeneous devices, which significantly deteriorates FL training efficiency. Meanwhile, frequent local NN training and transmission impose high energy consumption pressure on users. To tackle these issues, this paper proposes a premium multi-width NN-assisted hierarchical FL (HFL) framework in heterogeneous cloud-edge-device computing to achieve remarkable training speedup and energy conservation. Specifically, a heterogeneity-aware NN width coefficient determination algorithm, which flexibly assigns a subnet with a suitable width to each user device based on its computing ability, is first applied to shorten the HFL training latency. Subsequently, to integrate subnets with different width topologies, we design a width-aware adaptive NN model integration approach to effectively ensure the accuracy of the integrated global NN model. Finally, a latency-aware energy saving strategy is introduced to reduce energy consumption. Experimental results demonstrate that our proposed framework outperforms state-of-the-art benchmarks, and attains up to 42.42% enhancement in accuracy, 81.5% reduction in training latency, and 40.9% optimization in energy cost.

Jiazhen Chen, Zheng Ma 0011, Sichao Fu, Mingbin Feng, Tony S. Wirjanto, Weihua Ou

Graphs play a pivotal role in multimedia applications by integrating information to model complex relationships. Recently, graph class-incremental learning (GCIL) has garnered attention, allowing graph neural networks (GNNs) to adapt to evolving graph analytical tasks by incrementally learning new class knowledge while retaining knowledge of old classes. Existing GCIL methods primarily focus on a closed-set assumption, where all test samples are presumed to belong to previously known classes. Such assumption restricts their applicability in real-world scenarios, where unknown classes naturally emerge during inference, and are absent during training. In this paper, we explore a more challenging open-set graph class-incremental learning scenario with two intertwined challenges: catastrophic forgetting of old classes, which impairs the detection of unknown classes, and inadequate open-set recognition, which destabilizes the retention of learned knowledge. To address the above problems, a novel OGCIL framework is proposed, which utilizes pseudo-sample embedding generation to effectively mitigate catastrophic forgetting and enable robust detection of unknown classes. To be specific, a prototypical conditional variational autoencoder is designed to synthesize node embeddings for old classes, enabling knowledge replay without storing raw graph data. To handle unknown classes, we employ a mixing-based strategy to generate out-of-distribution (OOD) samples from pseudo in-distribution and current node embeddings. A novel prototypical hypersphere classification loss is further proposed, which anchors in-distribution embeddings to their respective class prototypes, while repelling OOD embeddings away. Instead of assigning all unknown samples into one cluster, our proposed objective function explicitly models them as outliers through prototype-aware rejection regions, ensuring a robust open-set recognition. Extensive experiments on five benchmarks demonstrate the effectiveness of OGCIL over existing GCIL and open-set GNN methods.

Xinhai Yan, Libing Wu, Zhuangzhuang Zhang, Bingyi Liu, Lijuan Huo, Jing Wang 0036

Federated Learning (FL) enables collaborative model training while preserving data privacy, but it is highly vulnerable to backdoor attacks. Most existing defense methods in FL have limited effectiveness due to their neglect of the model's over-reliance on backdoor triggers, particularly as the proportion of malicious clients increases. In this paper, we propose FedBAP, a novel defense framework for mitigating backdoor attacks in FL by reducing the model's reliance on backdoor triggers. Specifically, first, we propose a perturbed trigger generation mechanism that creates perturbation triggers precisely matching backdoor triggers in location and size, ensuring strong influence on model outputs. Second, we utilize these perturbation triggers to generate benign adversarial perturbations that disrupt the model's dependence on backdoor triggers while forcing it to learn more robust decision boundaries. Finally, we design an adaptive scaling mechanism to dynamically adjust perturbation intensity, effectively balancing defense strength and model performance. The experimental results demonstrate that FedBAP reduces the attack success rates by 0.22%-5.34%, 0.48%-6.34%, and 97.22%-97.6% under three types of backdoor attacks, respectively. In particular, FedBAP demonstrates outstanding performance against novel backdoor attacks.

Jiacheng Jiang, Yuan Meng, Chen Tang, Han Yu 0009, Qun Li, Zhi Wang 0001, Wenwu Zhu 0001

Current quantization-aware training (QAT) methods primarily focus on enhancing the performance of quantized models on in-distribution (I.D) data, while overlooking the potential performance degradation on out-of-distribution (OOD) data. In this paper, we first substantiate this problem through rigorous experiment, showing that QAT can lead to a significant OOD generalization performance degradation. Further, we find the contradiction between the perspective that flatness of loss landscape gives rise to superior OOD generalization and the phenomenon that QAT lead to a sharp loss landscape, can cause the above problem. Therefore, we propose a flatness-oriented QAT method, FQAT, to achieve generalizable QAT. Specifically, i) FQAT introduces a layer-wise freezing mechanism to mitigate the gradient conflict issue between dual optimization objectives (i.e., vanilla QAT and flatness). ii) FQAT proposes an disorder-guided adaptive freezing algorithm to dynamically determines which layers to freeze at each training step, effectively addressing the challenges caused by interference between layers. A gradient disorder metric is designed to help the algorithm identify unstable layers during training. Extensive experiments on influential OOD benchmark demonstrate the superiority of our method over state-of-the-art baselines under both I.D and OOD image classification tasks.

Lehao Lin, Baohua Fang, Ziheng Sun, Ke Wang 0013, Hong Kang, Wei Cai 0002

Current 3D model Level of Detail (LOD) methods require multiple models with varying detail levels to reduce client computational load, transmitting different models based on the user's distance to the object. However, this process consumes excessive network bandwidth and strains the client's memory and storage. To address this, we propose BS3 (Bézier Slicing for 3Ds), a middleware-enabled method that slices 3D meshes and fits the contours using Bézier curves. Acting as an intermediate layer, the BS3 middleware handles slicing, vectorization, sampling and reconstruction, allowing .bs3 files to be streamed only once and adjusted dynamically at different sampling rates. Our experiments demonstrate the efficiency and performance analysis of BS3, which shows that it can reduce network and storage burdens while keeping the display effect. We believe that BS3 will enhance 3D multimedia in the game, exhibition, digital museum, cultural heritage, metaverse, etc.

Xinbiao Gan, Qiang Zhang 0053, Tiejun Li, Chunye Gong, Kai Lu 0001

Graph has recently enabled substantial advances to the Web. Processing worldwide graphs with millions to billions, even trillions of edges in large-scale high-performance systems is pressing, but current graph processing engines are designed for small-scale graph processing beyond a few tens of computing nodes and are unable to scale well to large parallel systems because they are oblivious to imbalanced communication across the communication grid. Therefore, we present GraphWorld, a better approach to optimizing graph search in large parallel systems for world-wide web crawling and indexing.GraphWorld (i) features a new graph partitioning method to achieve better load balancing and minimize communication overhead across the row and column directions; (ii) designs an efficient hardware prefetching and caching mechanism that can gather, traverse, and scatter pipeline vertices to accelerate graph processing; and (iii) proposes υBFS: vectorization-based BFS for leveraging vectorization units equipped in modern high performance processors to further improve graph search.In addition, we used real-world graphs and benchmarks to demonstrate the effectiveness of GraphWorld. In particular, the GraphWorld-based Graph 500 tests on the Tianhe supercomputer are superior to the fastest systems in the latest Graph 500 lists. We finally apply GraphWorld to real-life graphs for the worldwide search of the Web, which outperforms the state-of-the-art graph partitioning and graph system by orders of magnitude.

Youbo Mao, Ziyang Kang, Pengfei Li, Jiyao Chen, Zenglin Yang, Zhijun Li 0002

With the increasing popularity of image and video analysis on mobile devices, high-throughput image inference has become essential. However, current mobile deep learning frameworks face key bottlenecks: high computational load in JPEG image recognition and low processor efficiency, which limit overall image processing throughput. To address these issues, this paper proposes the FCG framework (Frequency Domain model for CPU and GPU), a mobile JPEG inference framework based on frequency domain data and a hybrid parallel architecture that enables high-throughput inference for JPEG-encoded images on mobile devices. FCG decouples JPEG decoding from model inference by discarding the traditional RGB decoding process and retaining only the Huffman decoding. This decoding step is further accelerated through multi-core processing, significantly reducing the computational burden and latency during preprocessing. In light of the characteristics of frequency domain data and the heterogeneous CPU/GPU processors on mobile devices, FCG reconstructs the deep learning model to ensure recognition accuracy while optimizing resource utilization. By effectively allocating tasks and combining parallel and sequential execution, FCG optimizes processor resource utilization to achieve high throughput and low latency. FCG outperforms the state-of-the-art NN-Stretch by reducing latency by 36%. It also achieves significant throughput improvements—3.6x, 3.3x, and 2.8x—on CPU, GPU, and CPU+GPU configurations, respectively, compared to sequential inference systems. Additionally, FCG reduces power consumption by 56%, 35%, and 43% in these configurations.

Weiwu Pang, Rajrup Ghosh, Jiawei Yang 0006, Ziyu Wei, Branden Leong, Yue Wang, Ramesh Govindan

Outdoor AR applications on mobile devices need accurate estimates for the pose of the device. In this paper, we develop SplatPose, a novel pose estimation technique that uses a data-driven 3D modeling technique called Gaussian Splatting. SplatPose uses a trained Gaussian Splatting model to render an image at an estimated device location, then matches features with the camera image to estimate pose. % Because this matching can be fast, SplatPose can, in theory, estimate pose entirely on a mobile device, while existing approaches cannot. To this end, SplatPose trains Gaussian Splatting models to be robust to appearance changes, thereby improving accuracy. It also incorporates a novel fast renderer to improve rendering speed. Using an AR pose estimation benchmark dataset, we show that SplatPose outperforms the state-of-the-art in terms of accuracy, and is up to an order of magnitude faster on a mobile device.

Kewei Zhao, Xiaowei Hu 0001, Qinya Li

Unknown object detection aims to build detectors capable of identifying out-of-distribution objects, a critical need for applications like autonomous driving and traffic monitoring. However, limited device resources restrict existing methods from achieving accurate detection on the device side. Addressing this gap, this paper introduces a device-cloud collaborative framework named DCCUOD that enhances device model performance through efficient cloud collaboration. Our framework employs an energy-based sampling function on devices to target samples with unknown objects, coupled with a collaborative pseudo-labeling strategy to generate accurate pseudo-labels. Additionally, a two-stage training paradigm enables continuous improvements of device models on both known and unknown objects. Our study is the first to explore device-cloud collaborative learning for UOD tasks. Experimental results show that the device model is three times smaller and seven times faster than cloud models, with minimal performance trade-offs.

Dong Li, Yichen Niu, Ying Ai, Xiang Zou, Biqing Qi, Jianxing Liu

Large language models (LLMs) have demonstrated strong performance in natural language generation but remain limited in knowle- dge-intensive tasks due to outdated or incomplete internal knowledge. Retrieval-Augmented Generation (RAG) addresses this by incorporating external retrieval, with GraphRAG further enhancing performance through structured knowledge graphs and multi-hop reasoning. However, existing GraphRAG methods largely ignore the temporal dynamics of knowledge, leading to issues such as temporal ambiguity, time-insensitive retrieval, and semantic redundancy. To overcome these limitations, we propose Temporal GraphRAG (T-GRAG), a dynamic, temporally-aware RAG framework that models the evolution of knowledge over time. T-GRAG consists of five key components: (1) a Temporal Knowledge Graph Generator that creates time-stamped, evolving graph structures; (2) a Temporal Query Decomposition mechanism that breaks complex temporal queries into manageable sub-queries; (3) a Three-layer Interactive Retriever that progressively filters and refines retrieval across temporal subgraphs; (4) a Source Text Extractor to mitigate noise; and (5) a LLM-based Generator that synthesizes contextually and temporally accurate responses. We also introduce Time-LongQA, a novel benchmark dataset based on real-world corporate annual reports, designed to test temporal reasoning across evolving knowledge. Extensive experiments show that T-GRAG significantly outperforms prior RAG and GraphRAG baselines in both retrieval accuracy and response relevance under temporal constraints, highlighting the necessity of modeling knowledge evolution for robust long-text question answering. Our code is publicly available on the T-GRAG https://github.com/Arvin0313/T-GRAG.git

Adhi Widagdo, Teemu Kämäräinen, Ahmad Alhilal, Matti Siekkinen, Cheng-Hsin Hsu

Remote rendering enables high-fidelity virtual reality (VR) experiences on standalone headsets by offloading intensive graphics workloads to remote servers. However, streaming high-quality VR graphics imposes substantial bandwidth and latency challenges. Spatial compression is a form of foveation which addresses this challenge by leveraging the human visual system's varying acuity, allocating higher visual quality around the user's gaze while reducing resolution in the periphery. In this work, we implement three gaze-adaptive foveation methods: Dynamic Axis-Aligned Distortion Transmission (D-AADT2 and D-AADT3) and Dynamic Foveated Radial Warp (D-FRW)) of which only D-AADT2 has been previously presented. These methods dynamically adapt spatial compression based on gaze-tracking input, ensuring optimal perceptual quality. We integrate these methods together with their static counterparts into the open-source Air Light VR (ALVR) remote-rendering framework, enabling native (72 FPS) framerates. We conclude a comprehensive objective evaluation across diverse VR games and demonstrate that the dynamic methods significantly outperform traditional static approaches in both encoding efficiency and perceptual quality metrics. A complementary subjective user study further validates these findings, confirming that dynamic gaze-adaptive foveation substantially enhances visual quality, immersion, and user interaction experience.

Yang Zhao 0002, Shusheng Li, Xueshang Feng

As the development of lightweight deep learning algorithms, various deep neural network (DNN) models have been proposed for the remote sensing scene classification (RSSC) application. However, it is still challenging for these RSSC models to achieve optimal performance among model accuracy, inference latency, and energy consumption on resource-constrained edge devices. In this paper, we propose a lightweight RSSC framework, which includes a distilled global filter network (GFNet) model and an early-exit mechanism designed for edge devices to achieve state-of-the-art performance. Specifically, we first apply frequency domain distillation on the GFNet model to reduce model size. Then we design a dynamic early-exit model tailored for DNN models on edge devices to further improve model inference efficiency. We evaluate our E3C model on three edge devices across four datasets. Extensive experimental results show that it achieves an average of 1.3x speedup on model inference and over 40% improvement on energy efficiency, while maintaining high classification accuracy.

Fangxin Liu, Junjie Wang, Ning Yang 0012, Zongwu Wang, Junping Zhao, Li Jiang 0002, Haibing Guan

Transformer-based models have demonstrated remarkable performance in computer vision tasks. However, their increasing model size leads to substantial memory demands and higher latency, hindering practical deployment. This paper presents an adaptive dynamic layer-skipping framework based on Markov Decision Process, which determines optimal computational paths based on the current state of input samples. We introduce a Temporal Importance Difference Reward mechanism to address the credit assignment problem in layer-skipping decisions, and develop a knowledge distillation strategy using learnable cognitive tokens to compensate for information loss. Experiments on various models demonstrate that our method significantly reduces computational costs while maintaining accuracy, offering a practical solution for deploying high-performance Transformer models in resource-constrained environments. The code is available at https://github.com/wjjkhl/ASTER

Nan He, Yiming Chen, Zheng Jiang 0006, Song Yang 0002, Lifeng Sun

Federated Learning (FL) has become a powerful technique for collaborative model training across decentralized entities while preserving data privacy. Despite its potential, FL faces significant challenges, including communication overhead, resource heterogeneity, and data heterogeneity. Existing solutions fall short in addressing disparities in client resources and the errors introduced by direct model aggregation across heterogeneous clients. To tackle these issues, we propose DynFed, a novel federated learning framework that incorporates dynamic quantization bit-width allocation and multi-teacher knowledge distillation for model aggregation. DynFed dynamically adjusts quantization bit-widths to clients based on their resource heterogeneity, adapting these allocations according to variations in the local loss function during training. This adaptive quantization strategy optimizes resource utilization while preserving model performance. For model aggregation, DynFed utilizes a dynamic multi-teacher knowledge distillation approach, assigning the most suitable teacher model to each data sample based on a comprehensive evaluation score, thereby ensuring effective knowledge transfer even in the presence of quantization-induced errors. This method not only mitigates the negative effects of heterogeneous bit-widths but also leverages client model diversity to enhance the robustness of the global model. Extensive experimental results demonstrate the superiority of DynFed over state-of-the-art methods.

Zhe Sun, Qiang Xu, Qi Zhang 0029, Shan Liu 0001, Ge Li 0002

Compressing attributes of 3D point clouds remains challenging due to their inherent sparsity and irregular distribution. To address this, we propose an efficient framework based on sparse hierarchical Implicit Neural Representations (INRs). Specifically, we introduce a novel vertex-based INR framework, which integrates interpolation to enable accurate and compact implicit representations of point cloud attributes. To effectively capture the varying importance of latent features, we design an adaptive quantization scheme. Furthermore, we develop efficient level-wise entropy models to exploit dependencies within and across hierarchical levels. Finally, point cloud attributes are reconstructed from concatenated multi-resolution latent representations via a sparse convolution-based reconstruction module. Experimental results demonstrate that our approach significantly outperforms previous INR-based methods, achieving superior performance compared to the latest G-PCC (TMC13v28) standard and state-of-the-art learning-based methods.