论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 46 / 81 页

Cunhang Fan, Sheng Zhang, Jingjing Zhang, Enrui Liu, Xinhui Li, Gangming Zhao, Zhao Lv

Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in the precise reconstruction of minute-level continuous imagined speech: traditional models struggle to balance the efficiency of temporal dependency modeling and information retention in long-sequence decoding. To address this issue, this paper proposes the Dynamic Multiscale Fusion Network (DMF2Mel), which consists of four core components: the Dynamic Contrastive Feature Aggregation Module (DC-FAM), the Hierarchical Attention-Guided Multi-Scale Network (HAMS-Net), the SplineMap attention mechanism, and the bidirectional state space module (convMamba). Specifically, the DC-FAM separates speech-related ''foreground features'' from noisy ''background features'' through local convolution and global attention mechanisms, effectively suppressing interference and enhancing the representation of transient signals. HAMS-Net, based on the U-Net framework, achieves cross-scale fusion of high-level semantics and low-level details. The SplineMap attention mechanism integrates the Adaptive Gated Kolmogorov-Arnold Network (AGKAN) to combine global context modeling with spline-based local fitting. The convMamba captures long-range temporal dependencies with linear complexity and enhances nonlinear dynamic modeling capabilities. Results on the SparrKULee dataset show that DMF2Mel achieves a Pearson correlation coefficient of 0.074 in mel spectrogram reconstruction for known subjects (a 48% improvement over the baseline) and 0.048 for unknown subjects (a 35% improvement over the baseline).Code is available at: https://github.com/fchest/DMF2Mel.

Yang Hu, Jingui Ma, Yucheng Yang, Jie Liang, Jinbo Yan, Jiahao Wu, Jiayu Yang, Yang Deng, Ronggang Wang

3D Gaussian Splatting (3DGS) has emerged as a promising framework for real-time radiance field rendering due to its high fidelity and explicit scene modeling. However, its practical deployment in the multimedia domain remains limited by excessive memory usage stemming from redundant and memory-inefficient Gaussian primitives. In this paper, we propose SOC-GS, a novel compression framework that enhances the anchor-based 3DGS representation through perceptually guided and structural optimization. Specifically, we begin by introducing the Perceptual Relevance Score (PRS), with a Gumbel noise perturbation applied to facilitate sparse Top-K selection of Gaussians critical for densification, significantly reducing the number of anchors. Further, we stabilize training and prevent premature overfitting the high-frequency noise using a Joint Resolution-Blur Training strategy, with guidance from Total Variation Loss, enabling coarse-to-fine learning with the consistency of spatial distribution throughout training. Finally, a Spatial Condition-based Prediction module is employed to further reduce storage while preserving comparable quality. Extensive experiments on three benchmark datasets demonstrate that our method achieves an average of 34% reduction in model size when compared to existing state-of-the-art compression method (126 × compression on vanilla 3DGS), while maintaining comparable--or even superior--rendering quality.

Yixuan Gao, Xiongkuo Min, Jinliang Han, Yuqin Cao, Sijing Wu, Yunze Dou, Guangtao Zhai

With the rise of Text-to-Image (T2I) models, generating face images from text prompts has emerged as a prominent research area. However, evaluating the quality of these generated face images, particularly with respect to fine-grained facial attributes, remains a significant challenge. To address this, we introduce the Fine -grained Text-to-Face Image Quality Assessment (FineTFIQA) database, which is designed to evaluate the ability of T2I models to generate fine-grained face images. To the best of our knowledge, this database is the largest of its kind, containing 7,218 face images generated from 1,000 text prompts that cover 111 distinct facial attributes. A large group of subjects was invited to assess the quality of text-to-face images on four evaluation dimensions: perceptual quality, human likeness, attractiveness, and consistency. Additionally, we develop the Multi-Dimensional Text-to-Face Image Quality Assessment (MDTFIQA) method based on the Large Language Model (LLM), which combines both face image features and text features to evaluate generated images on all evaluation dimensions. Extensive experimental results demonstrate that traditional face image assessment methods and general image quality assessment methods are inadequate for accurately evaluating generated text-to-face images. Our method significantly outperforms these existing methods on all evaluation dimensions, proving to be an effective method for assessing the quality of generated text-to-face images.

De Li, Zhou Tan, Qiyu Li, Zeming Gan, Tiange Xia, Jinyan Wang, Xianxian Li

Federated graph classification has emerged as a promising paradigm for privacy-preserving graph learning across distributed clients. However, real-world federated scenarios often suffer from severe data heterogeneity and label noise, which significantly degrade model performance. To address these challenges, we propose FedRog, a robust and personalized federated graph neural network framework that improves generalization under non-IID and noisy label settings. FedRog introduces a parameter-aware selection and fine-tuning mechanism to align global and local representations, and a neighbor embedding consistency constraint to enhance robustness against noisy supervision. Furthermore, a fine-grained, importance-guided global aggregation strategy based on Fisher information is employed to mitigate unreliable updates from low-quality clients. We conduct extensive experiments on 16 graph classification datasets under five heterogeneous data partition settings. Results show that FedRog consistently achieves competitive or superior performance compared to 14 baselines in terms of both accuracy and robustness under clean and noisy conditions.

Sijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao, Huiyu Duan, Wei Sun 0029, Guangtao Zhai

Face video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to human faces. However, FVQA is rarely explored due to the lack of large-scale FVQA datasets. To fill this gap, we present the first large-scale in-the-wild FVQA dataset, FVQ-20K, which contains 20,000 in-the-wild face videos together with corresponding mean opinion score (MOS) annotations. Along with the FVQ-20K dataset, we further propose a specialized FVQA method named FVQ-Rater to achieve human-like rating and scoring for face video, which is the first attempt to explore the potential of large multimodal models (LMMs) for the FVQA task. Concretely, we elaborately extract multi-dimensional features including spatial features, temporal features, and face-specific features (i.e., portrait features and face embeddings) to provide comprehensive visual information, and take advantage of the LoRA-based instruction tuning technique to achieve quality-specific fine-tuning, which shows superior performance on both FVQ-20K and CFVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the FVQ-20K dataset and FVQ-Rater method in promoting the development of FVQA. The code and dataset will be released at: https://github.com/wsj-sjtu/FVQ.

Gyeongjin Kim, Sebin Lee, Daye Kim, Jungjin Lee, Minju Kim

Live-streamed concerts have become a new cultural phenomenon, yet they struggle to replicate the collective emotional experience of in-person events. Traditional text-based chats often cause information overload and distraction, diminishing the sense of shared experience. To address this challenge, we present VibeOn, a multimodal interaction system designed to foster collective engagement and socio-emotional connection among remote audiences. Developed through a formative study and an iterative design process, VibeOn integrates features such as chat and emoji recommendations, chat highlights, avatar-based cheering, ambient visualization, and concert-specific layout. A user study with 40 participants demonstrated that VibeOn significantly enhanced social connectedness, sense of community, and collective effervescence compared to a conventional chat interface, while maintaining high usability. Our findings indicate that VibeOn enables audiences to feel a shared, extraordinary experience beyond simply watching, highlighting its potential to enrich collective emotions in large-scale online events.

Zitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma, Jiarui Wang, Liu Yang, Shiqi Gao, Xiaoyu Wang, Jia Wang 0004, Xiongkuo Min 等

The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.

Lei Chen 0093

Super-resolution (SR) images, generated by advanced algorithms to enhance resolution under hardware constraints, are increasingly applied across various multimedia tasks. However, the absence of paired high-resolution (HR) reference image and the inherent ill-posedness of SR reconstruction present key challenges for SR image quality assessment (SR-IQA). Full-reference methods become inapplicable, while the reduced-reference methods relying on one low-resolution (LR) image offer limited reliability. To address these issues, I propose the SQer, a no-reference SR-IQA method based on a graph perceptron with semantic fidelity. The SQer first extracts perceptual and hierarchical SR image features using a superposition nonlinear feature pooling. These features are transformed into graph vector representations, allowing semantic information learning via a graph-structured attention perceptron. Finally, the resulting graphs are globally average-pooled into a semantic embedding, which is then processed by a multilayer perceptron to predict the SR image quality score. Extensive experiments on multiple SR-IQA benchmarks demonstrate that my proposed SQer significantly outperforms existing state-of-the-art reference-based methods, exhibiting superior accuracy and a stronger ability to capture fine-grained perceptual cues and SR-specific artifacts. The SQer method provides a promising direction for guiding the optimization and application of image super-resolution models.

Yongyang Zhou, Fanglue Zhang, Zichen Wang, Lei Zhang

3D Gaussian Splatting (3DGS) has demonstrated impressive capabilities in novel view synthesis. However, rendering reflective objects remains a significant challenge, particularly in inverse rendering and relighting. We introduce RTR-GS, a novel inverse rendering framework capable of robustly rendering objects with arbitrary reflectance properties, decomposing BRDF and lighting, and delivering credible relighting results. Given a collection of multi-view images, our method effectively recovers geometric structure through a hybrid rendering model that combines forward rendering for radiance transfer with deferred rendering for reflections. This approach successfully separates high-frequency and low-frequency appearances, mitigating floating artifacts caused by spherical harmonic overfitting when handling high-frequency details. We further refine BRDF and lighting decomposition using an additional physically-based deferred rendering branch. Experimental results show that our method enhances novel view synthesis, normal estimation, decomposition, and relighting while maintaining efficient training inference process.

Weizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou 0004, Xiaoxuan Tang, Jiajun Bu, Yong Li 0004, Wei Jiang 0041

Graphical User Interface (GUI) agents possess significant commercial and social value, and GUI agents powered by advanced multimodal large language models (MLLMs) have demonstrated remarkable potential. Currently, existing GUI agents usually utilize sequential episodes of multi-step operations across pages as the prior GUI knowledge, which fails to capture the complex transition relationship between pages, making it challenging for the agents to deeply perceive the GUI environment and generalize to new scenarios. Therefore, we design an automated pipeline to transform the sequential episodes into page graphs, which explicitly model the graph structure of the pages that are naturally connected by actions. To fully utilize the page graphs, we further introduce Retrieval-Augmented Generation (RAG) technology to effectively retrieve reliable perception guidelines of GUI from them, and a tailored multi-agent framework PG-Agent with task decomposition strategy is proposed to be injected with the guidelines so that it can generalize to unseen scenarios. Extensive experiments on various benchmarks demonstrate the effectiveness of PG-Agent, even with limited episodes for page graph construction. Our codes will be publicly available at https://github.com/chenwz-123/PG-Agent.

Xuan Zhang, Sin Chee Chin, Jing-Hao Xue, Xiaochen Yang, Wenming Yang

Long-tailed out-of-distribution learning aims to reduce performance bias in long-tailed in-distribution (ID) data while rejecting out-of-distribution (OOD) samples, which are often mistaken for under-represented tail classes. To achieve OOD detection, existing methods incorporate an outlier exposure (OE) term into the long-tailed recognition (LTR) loss. However, as we prove in this paper, the OE term induces a gradient conflict with the ID objectives, especially for tail classes, thereby contradicting the core motivation of LTR. To avoid the ID-OOD dilemma, we propose Dynamic Ambiguity-aware Recalibration for Logits (DARL), an ambiguity-guided long-tailed OOD learning approach, grounded on two theoretical insights. First, we show that the mixed ID data can mitigate the conflict in OE training and exhibits higher intrinsic ambiguity than the original ID data, thus able to serve as a surrogate for real OOD data. Second, we introduce an ambiguity-aware logit adjustment that can dynamically calibrate the class margins using energy-based ambiguity metrics, effectively reducing early-stage bias while avoiding late-stage overfitting. Extensive experiments show that DARL achieves the overall state-of-the-art performance of long-tailed OOD learning. Moreover, compared with the OE methods, DARL trains solely on the ID data, which can reduce the data requirements by 80%. The code is available in https://github.com/XuanZhang-A/DARL.

Yitong Zhu, Zhuowen Liang, Yiming Wu, Tangyao Li, Yuyang Wang 0002

Cybersickness remains a major obstacle to the widespread adoption of immersive virtual reality (VR), particularly in consumer-grade environments. While prior methods rely on invasive signals such as electroencephalography (EEG) for high predictive accuracy, these approaches require specialized hardware and are impractical for real-world applications. In this work, we propose a scalable, deployable framework for personalized cybersickness prediction leveraging only non-invasive signals readily available from commercial VR headsets, including head motion, eye tracking, and physiological responses. Our model employs a modality-specific graph neural network enhanced with a Difference Attention Module to extract temporal-spatial embeddings capturing dynamic changes across modalities. A cross-modal alignment module jointly trains the video encoder to learn personalized traits by aligning video features with sensor-derived representations. Consequently, the model accurately predicts individual cybersickness using only video input during inference. Experimental results show our model achieves 88.4% accuracy, closely matching EEG-based approaches (89.16%), while reducing deployment complexity. With an average inference latency of 90ms, our framework supports real-time applications, ideal for integration into consumer-grade VR platforms without compromising personalization or performance. The code will be relesed at https://github.com/U235-Aurora/PTGNN.

Songpei Xu, Xuri Ge, Chaitanya Kaul, Roderick Murray-Smith

This study aims to utilise mid-air hand-pose movements to implement various interactive controls, e.g. dial and slider controlling, through independent low-dimensional embeddings. Towards this, we develop a novel adjustable hand-pose space disentanglement approach for a learnable VAE-based high-to-low dimensional embedding model ( HandSolo ). It disentangles the latent embeddings into multiple independent one- or two-dimensional embedding spaces, enabling independent control. HandSolo allows multi-dimensional settings and multi-DOF combinations, providing a new paradigm for flexible and extensible hand-pose interaction systems. Additionally, to exploit model potential and make user interaction comfortable, we propose a visual interaction evaluation strategy (VIEs) to help system designers understand model capability and user habits. Finally, we provide an example virtual interaction system that integrates various virtual interaction objects, showing how our innovations improve their interaction capabilities. Experimental user studies demonstrate the effectiveness of our embedding-disentanglement designs, including discovery experiment (n=4) for VIEs, inspiration experiment (n=4) for approach extensibility, and exploration experiment (n=8) for the virtual interaction system.

Sihan Zhao, Zixuan Wang 0026, Tianyu Luan, Jia Jia 0001, Wentao Zhu 0004, Jiebo Luo 0001, Junsong Yuan 0001, Nan Xi

Human motion generation has found widespread applications in AR/VR, film, sports, and medical rehabilitation, offering a cost-effective alternative to traditional motion capture systems. However, evaluating the fidelity of such generated motions is a crucial, multifaceted task. Although previous approaches have attempted at motion fidelity evaluation using human perception or physical constraints, there remains an inherent gap between human-perceived fidelity and physical feasibility. Moreover, the subjective and coarse binary labeling of human perception further undermines the development of a robust data-driven metric. We address these issues by introducing a physical labeling method. This method evaluates motion fidelity by calculating the minimum modifications needed for a motion to align with physical laws. With this approach, we are able to produce fine-grained, continuous physical alignment annotations that serve as objective ground truth. With these annotations, we propose PP-Motion, a novel data-driven metric to evaluate both physical and perceptual fidelity of human motion. To effectively capture underlying physical priors, we employ Pearson's correlation loss for the training of our metric. Additionally, by incorporating a human-based perceptual fidelity loss, our metric can capture fidelity that simultaneously considers both human perception and physical alignment. Experimental results demonstrate that our metric, PP-Motion, not only aligns with physical laws but also aligns better with human perception of motion fidelity than previous work.

Xiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen 0003, Tong Zhu 0003, Leida Li

Aesthetic Image Cropping (AIC) aims to improve the visual appeal of images by removing redundant content while preserving attractive elements. Despite the encouraging progresses achieved in data-driven approaches, most existing models struggle to understand user intentions, particularly for diversified scenes with multiple subjects. Moreover, they can only provide cropping results without explanations, which further restricts their usability in real-world applications. Motivated by the above facts, we introduce InstructCrop : a multimodal large language model (MLLM)-based AIC framework, which can understand user instructions and provide explanatory reasons for cropping results. Specifically, we first build a multimodal Image Cropping Instruction Tuning (ICIT) dataset through a cost-effective paradigm by generating high-quality instruction tuning data based on the existing cropping datasets. Then, we embed dynamic domain knowledge into the cropping model by integrating cropping-aware experts of aesthetic assessment and composition classification. Finally, we adapt MLLMs to generate the cropping results and corresponding explanations. Quantitative and qualitative experiments on three benchmark datasets demonstrate that InstructCrop enables effective and interpretable image cropping, which aligns better with user intentions. Data and code are available at https://github.com/sxfly99/InstructCrop.

Junzhe Zhang, Chengfeng Han, Dandan Ding, Zhan Ma 0001

Point cloud compression (PCC) is indispensable for the upcoming holographic communication, enabling efficient transmission and real-time interaction with high-fidelity 3D data. As a mature international PCC standard, MPEG G-PCC holds promise for widespread applications due to its support for various point cloud types and ease of implementation. However, the geometry compression performance of G-PCC is limited, which severely impacts the user quality of experience (QoE). To address this, we propose GeoQE, an enhancement model that seamlessly integrates with the G-PCC decoder to mitigate compression artifacts and improve QoE. GeoQE introduces two key techniques: (1) a quantizer-guided expansion operation that adaptively handles distortions at varying levels, and (2) a spatiotemporal mechanism that leverages correlations within the current frame and across adjacent frames, allowing effective enhancement even with a lightweight network. Experiments show that GeoQE delivers state-of-the-art performance on both dense (e.g., VR/AR) and sparse (e.g., LiDAR for autonomous driving) point clouds. Operating at around 4 fps on a 3090Ti GPU with a compact 1.6 MB model, it achieves much lower computational complexity than existing methods, which is attractive for practical applications.

Richen Liu, Lingyu Sun, Xuefeng Huang, Yiran Li, Jiang Zhang 0002, Siru Chen, Zhouhao Wu, Ayush Kumar 0004, Chufan Lai

Interactive data illustrations in an immersive environment are challenging due to their inherent ambiguities during the interaction. These challenges are introduced by visual clutter and 3D occlusions resulting from depth information, as well as the relatively inefficient fine-grained manipulations required by handle controllers on immersive devices. In this paper, we propose Meta-Illustrator, an illustration transfer tool to generate immersive 3D illustrations for a volumetric data with the 2D illustrated results transferred from its one or multiple 2D slices (images). Initially, the slices can be illustrated by users expressively, owing to the plenty of the existing mature 2D sketching techniques and image processing algorithms. Then the 2D illustrated results on the slices can be intelligently transferred from their 2D image space to 3D volumetric space by Meta-Illustrator. Compared to the state-of-the-art image-to-image style transfer neural networks, which are either computation-intensive or memory-intensive, the proposed 2D-to-3D transferring approach can be built on a desktop PC without training. We demonstrate the usability, expressiveness, and effectiveness of Meta-Illustrator by both quantitative and qualitative evaluations.

Tong Liu, Zhiwei Fan, Guanyan Peng, Haodan Zhang, Yucheng Zhang, Zhen Wang 0071, Pengjin Xie, Liang Liu 0001

Short video streaming has become a dominant paradigm in digital media, characterized by rapid swiping interactions and diverse media content. A key technical challenge is designing an effective preloading strategy that dynamically selects and prioritizes download tasks from an evolving playlist, balancing Quality of Experience (QoE) and bandwidth efficiency under practical commercial constraints. However, real-world analysis reveals critical limitations of existing approaches: (1) insufficient adaptation of download task sizes to dynamic conditions, and (2) watch-time prediction models that are difficult to deploy reliably at scale. In this paper, we propose DeLoad, a novel preloading framework that addresses these issues by introducing dynamic task sizing and a practical, multi-dimensional watch-time estimation method. Additionally, a Deep Reinforcement Learning (DRL)-enhanced agent is trained to optimize the download range decisions adaptively. Extensive evaluations conducted on an offline testing platform, leveraging massive real-world network data, demonstrate that DeLoad achieves significant improvements in QoE metrics (34.4%-87.4% gain). Furthermore, after deployment on a large-scale commercial short-video platform, DeLoad has increased overall user watch-time by 0.9‰ while simultaneously reducing rebuffering events and 3.76% bandwidth consumption.

Wenhao Li, Xiu Su, Jingyi Wu, Feng Yang, Yang Liu 0246, Yi Chen, Shan You, Chang Xu 0002

Large Vision-Language Models (LVLMs) have demonstrated remarkable advancements in numerous areas such as multimedia. However, hallucination issues significantly limit their credibility and application potential. Existing mitigation methods typically rely on external tools or the comparison of multi-round inference, which significantly increase inference time. In this paper, we propose SElf-Evolving Distillation (SEED), which identifies hallucinations within the inner knowledge of LVLMs, isolates and purges them, and then distills the purified knowledge back into the model, enabling self-evolution. Furthermore, we identified that traditional distillation methods are prone to inducing void spaces in the output space of LVLMs. To address this issue, we propose a Mode-Seeking Evolving approach, which performs distillation to capture the dominant modes of the purified knowledge distribution, thereby avoiding the chaotic results that could emerge from void spaces. Moreover, we introduce a Hallucination Elimination Adapter, which corrects the dark knowledge of the original model by learning purified knowledge. Extensive experiments on multiple benchmarks validate the superiority of our SEED, demonstrating substantial improvements in mitigating hallucinations for representative LVLM models such as LLaVA-1.5 and InternVL2. Remarkably, the F1 score of LLaVA-1.5 on the hallucination evaluation metric POPE-Random improved from 81.3 to 88.3.