论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 19 / 81 页

Divya Kothandaraman, Ming Lin 0003, Dinesh Manocha

We introduce a novel approach for concept blending in pretrained text-to-image diffusion models, aiming to generate images at the intersection of multiple text prompts. At each time step during diffusion denoising, our algorithm forecasts predictions w.r.t. the generated image and makes informed text conditioning decisions. Central to our method is the unique analogy between diffusion models, which are rooted in non-equilibrium thermodynamics, and the Black-Scholes model for financial option pricing. By drawing parallels between key variables in both domains, we derive a robust algorithm for concept blending that capitalizes on the Markovian dynamics of the Black-Scholes framework. Our text-based concept blending algorithm is data-efficient, meaning it does not need additional training. Furthermore, it operates without human intervention or hyperparameter tuning. We highlight the benefits of our approach by comparing it qualitatively and quantitatively to other text based concept blending techniques, including linear interpolation, alternating prompts, step-wise prompt switching, and CLIP-guided prompt selection across various scenarios such as single object per text prompt, multiple objects per text prompt and backgrounds. Our work shows that financially inspired techniques can enhance text-to-image concept blending in generative AI, paving the way for broader innovation. Code is available at https://github.com/divyakraman/BlackScholesDiffusion2024.

Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan

Recent advances in NeRF and 3DGS have significantly enhanced the efficiency and quality of 3D content synthesis. However, efficient personalization of generated 3D content remains a critical challenge. Current 3D personalization approaches predominantly rely on knowledge distillation-based methods, which require computationally expensive retraining procedures. To address this challenge, we propose Invert3D, a novel framework for convenient 3D content personalization. Nowadays, vision-language models such as CLIP enable direct image personalization through aligned vision-text embedding spaces. However, the inherent structural differences between 3D content and 2D images preclude direct application of these techniques to 3D personalization. Our approach bridges this gap by establishing alignment between 3D representations and text embedding spaces. Specifically, we develop a camera-conditioned 3D-to-text inverse mechanism that projects 3D contents into a 3D embedding aligned with text embeddings. This alignment enables efficient manipulation and personalization of 3D content through natural language prompts, eliminating the need for computationally retraining procedures. Extensive experiments demonstrate that Invert3D achieves effective personalization of 3D content.

Xiufeng Huang, Ziyuan Luo, Qi Song 0003, Ruofei Wang, Renjie Wan

The growing popularity of 3D Gaussian Splatting (3DGS) has intensified the need for effective copyright protection. Current 3DGS watermarking methods rely on computationally expensive fine-tuning procedures for each predefined message. We propose the first generalizable watermarking framework that enables efficient protection of Splatter Image-based 3DGS models through a single forward pass. We introduce GaussianBridge that transforms unstructured 3D Gaussians into Splatter Image format, enabling direct neural processing for arbitrary message embedding. To ensure imperceptibility, we design a Gaussian-Uncertainty-Perceptual heatmap prediction strategy for preserving visual quality. For robust message recovery, we develop a dense segmentation-based extraction mechanism that maintains reliable extraction even when watermarked objects occupy minimal regions in rendered views. Project page: https://kevinhuangxf.github.io/marksplatter.

Xiaohao Liu, Xiaobo Xia, Zhuo Huang, See-Kiong Ng, Tat-Seng Chua

Multi-modal learning has achieved remarkable success by integrating information from various modalities, achieving superior performance in tasks like recognition and retrieval compared to uni-modal approaches. However, real-world scenarios often present novel modalities that are unseen during training due to resource and privacy constraints, a challenge current methods struggle to address. This paper introduces Modality Generalization (MG), which focuses on enabling models to generalize to unseen modalities. We define two cases: Weak MG, where both seen and unseen modalities can be mapped into a joint embedding space via existing perceptors, and Strong MG, where no such mappings exist. To facilitate progress, we propose a comprehensive benchmark featuring multi-modal algorithms and adapt existing methods that focus on generalization. Extensive experiments highlight the complexity of MG, exposing the limitations of existing methods and identifying key directions for future research. Our work provides a foundation for advancing robust and adaptable multi-modal models, enabling them to handle unseen modalities in realistic scenarios.

Esen K. Tütüncü, Lissette Lemus, Kris Pilcher, Holger Sprengel, Jordi Sabater-Mir

Commonaiverse is an interactive installation exploring human emotions through full-body motion tracking and real-time AI feedback. Participants engage in three phases: Teaching, Exploration and the Cosmos Phase, collaboratively expressing and interpreting emotions with the system. The installation integrates MoveNet for precise motion tracking and a multi-recommender AI system to analyze emotional states dynamically, responding with adaptive audiovisual outputs. By shifting from top-down emotion classification to participant-driven, culturally diverse definitions, we highlight new pathways for inclusive, ethical affective computing. We discuss how this collaborative, out-of-the-box approach pushes multimedia research beyond single-user facial analysis toward a more embodied, co-created paradigm of emotional AI. Furthermore, we reflect on how this reimagined framework fosters user agency, reduces bias, and opens avenues for advanced interactive applications.

Weilin Wu, Shifan Yang, Qizhao Lin, Xinghong Chen, Kunping Yang, Jing Wang, Guannan Chen

Low-light image enhancement (LLIE) aims to restore low-light images to normal lighting conditions by improving their illumination and fine details, thereby facilitating efficient execution of downstream visual tasks. Traditional LLIE methods improve image quality but often introduce high-frequency artifacts, which are difficult to eliminate, hindering detail recovery and quality enhancement in LLIE. To solve this problem, we introduce a novel perspective: instead of traditional artifact suppression, sparsification-induced artifacts are repurposed as constructive regularization signals to guide detail recovery. By analyzing the impact of sparsified frequency components and their role in reconstruction artifacts, a detailed mathematical framework is presented. Specifically, we propose a novel loss function SASW-Loss which combining Sparse Artifact Similarity Loss (SAS-Loss) and Walsh-Hadamard Coefficient Loss (WHC-Loss). SAS-Loss mitigates the over-compensation of missing frequencies, helping the network recover structural details, while WHC-Loss optimizes the frequency-domain representation, restoring luminance, suppressing noise, and enhancing both structure and details. Extensive experiments show that our approach outperforms existing state-of-the-art methods, achieving superior performance in structural detail preservation and noise suppression. These results validate the effectiveness of our new perspective, which leverages sparsification artifacts to guide detail recovery, demonstrating significant improvements and robust performance across multiple models, and opening new avenues for future research. The code is available at https://github.com/werringwu/SASW.git.

Matyas Bohacek, Ignacio Vilanova Echavarri

Generative Artificial Intelligence (GAI) has experienced exponential growth in recent years, partly facilitated by the abundance of large-scale open-source datasets. These datasets are often built using unrestricted and opaque data collection practices. While most literature focuses on the development and applications of GAI models, the ethical and legal considerations surrounding the creation of these datasets are often neglected. In addition, as datasets are shared, edited, and further reproduced online, information about their origin, legitimacy, and safety often gets lost. To address this gap, we introduce the Compliance Rating Scheme (CRS), a framework designed to evaluate dataset compliance with critical transparency, accountability, and security principles. We also release an open-source Python library built around data provenance technology to implement this framework, allowing for seamless integration into existing dataset-processing and AI training pipelines across multiple data modalities, including images, video, audio, and 3D assets. The library is simultaneously reactive and proactive, as in addition to evaluating the CRS of existing datasets, it equally informs responsible scraping and construction of new datasets.

Xinyu Zhang 0017, Dong Gong, Zicheng Duan, Anton van den Hengel, Lingqiao Liu

Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and musical beats enhances viewer engagement and visual appeal, particularly in music videos, promotional content, and cinematic editing. Existing methods typically depend on labor-intensive manual cutting, speed adjustments, or heuristic-based editing techniques to achieve synchronization. While some generative models handle joint video and music generation, they often entangle the two modalities, limiting flexibility in aligning video to music beats while preserving the full visual content. In this paper, we propose a novel and efficient framework-termed MVAA (Music-Video Auto-Alignment)-that automatically edits video to align with the rhythm of a given music track while preserving the original visual content. To enhance flexibility, we modularize the task into a two-step process in our MVAA: aligning motion keyframes with audio beats, followed by rhythm-aware video inpainting. Specifically, we first insert keyframes at timestamps aligned with musical beats, then use a frame-conditioned diffusion model to generate coherent intermediate frames, preserving the original video's semantic content. Since comprehensive test-time training can be time-consuming, we adopt a two-stage strategy: pretraining the inpainting module on a small video set to learn general motion priors, followed by rapid inference-time fine-tuning for video-specific adaptation. This hybrid approach enables adaptation within ~10 minutes with one epoch on a single NVIDIA 4090 GPU using CogVideoX-5b-I2V [77] as the backbone. Extensive experiments show that our approach can achieve high-quality beat alignment and visual smoothness. User studies further validate the natural rhythmic quality of the results, confirming their effectiveness for practical music-video editing. The code is available at: zhangxinyu-xyz.github.io/MVAA

Feida Liu, Yifan Wang, Jiaqi Zheng 0001, Boxi Liu, Guihai Chen

Low-latency interactive video streaming services critically depend on robust congestion control algorithms (CCA). However, existing CCAs often exhibit frequent self-induced oscillations from unconstrained probing, causing periodic heavy queuing and degrading the user experience. We propose Themis, a novel end-to-end CCA designed to achieve stabilized near-zero queuing delay with high bitrate. By precisely quantifying the trade-off between bitrate and queuing delay within a utility feedback mechanism, Themis effectively controls the amplitude and frequency of probing while ensuring fairness. In parallel, Themis incorporates the awareness of the queuing state through an adaptive-pacing method, combined with utility feedback to guide a three-phase bitrate adjustment strategy. This enables rapid and stable convergence to the optimal utility. We implemented Themis in QUIC, evaluated it on the Mahimahi and conducted a 90-day large-scale A/B test in a real-world network. Compared to state-of-the-art CCAs, Themis effectively suppresses self-induced oscillations, increases the average frame bitrate by 66.8%, and reduces the average frame delay by 13.5%, demonstrating a superior trade-off between high bitrate and low queuing delay.

Daoxu Sheng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao

Video streaming platforms and existing ABRs traditionally assume uninterrupted sequential playback, yet users frequently skip to points of interest-a fundamental mismatch causing degradation of quality of experience at high-interest segments while wasting bandwidth on skipped content. We address this through our Hotspot-Aware Joint Optimization framework, which reframes video streaming as a non-monotonic optimization problem with discontinuous state transitions caused by navigation events. Our framework jointly optimizes adaptive bitrate decisions and buffer management by leveraging viewer engagement patterns to predict navigation behavior. Our approach combines: (1) a mathematical formulation capturing state discontinuities in non-sequential viewing, (2) self-supervised models predicting navigation targets using only aggregate viewing data, and (3) hotspot-aware ABR and buffer management algorithms implemented through our Streaming Local Search (SLS) technique that dynamically prioritize quality for frequently-watched segments. Evaluation across diverse content and network conditions demonstrates our framework delivers 38.2% higher quality in hotspot regions, 32.5% reduced navigation delays, and 27.1% improved resource efficiency compared to traditional methods. These improvements establish a foundation for streaming systems that adapt to both network conditions and content structure, aligning resource allocation with actual viewing patterns.

Yufeng Chen 0007, Umakant Kulkarni, Voicu Popescu, Sonia Fahmy

Immersive virtual reality (VR) experiences require transmission and rendering of large-scale 3D content, often represented as point clouds or polygon meshes. Unfortunately, existing networked VR systems often fail to fully exploit the flexibility of VR data representations. To address this problem, we propose a cross-layer design that elevates a network data unit to a usable rendering unit for VR applications. Our aim is to bridge the gap between networks and applications in order to enhance visual quality, especially over constrained and variable networks. Our approach, Rendering Unit that is Network-aware (RUN), with two variants, RUN-Packet and RUN-Hybrid, includes mechanisms to effectively utilize network data units when encoding, transmitting, decoding, and rendering. Specifically, we develop additive detail refinement mechanisms and address streaming challenges such as head-of-line (HoL) blocking. We prototype our system in Unity 3D and evaluate it using synthetic network environments and real network traces. Our results with both static and dynamic point clouds demonstrate that RUN significantly reduces stalls and delivers smoother frame updates, enhancing visual quality.

Ruonan Chai, Yixiang Zhu, Xinjiao Li, Jiawei Li 0009, Zili Meng, Dirk Kutscher

Real-time streaming of point cloud video, characterized by massive data volumes and high sensitivity to packet loss, remains a key challenge for immersive applications under dynamic network conditions. While connection-oriented protocols such as TCP and more modern alternatives like QUIC alleviate some transport-layer inefficiencies, including head-of-line blocking, they still retain a coarse-grained, segment-based delivery model and a centralized control loop that limit fine-grained adaptation and effective caching. We introduce INDS (Incremental Named Data Streaming), an adaptive streaming framework based on Information-Centric Networking (ICN) that rethinks delivery for hierarchical, layered media. INDS leverages the Octree structure of point cloud video and expressive content naming to support progressive, partial retrieval of enhancement layers based on consumer bandwidth and decoding capability. By combining time-windows with Group-of-Frames (GoF), INDS's naming scheme supports fine-grained in-network caching and facilitates efficient multi-user data reuse. INDS can be deployed as an overlay, remaining compatible with QUIC-based transport infrastructure as well as future Media-over-QUIC (MoQ) architectures, without requiring changes to underlying IP networks. Our prototype implementation shows up to 80% lower delay, 15-50% higher throughput, and 20-30% increased cache hit rates compared to state-of-the-art DASH-style systems. Together, these results establish INDS as a scalable, cache-friendly solution for real-time point cloud streaming under variable and lossy conditions, while its compatibility with MoQ overlays further positions it as a practical, forward-compatible architecture for emerging immersive media systems.

Runjie Wang, Kemi Chen, Shuijie Li, Mingkai Chen 0001, Tiesong Zhao

Nowadays, haptic data has gained a fast-growing volume with enormous interaction points during human-computer interaction and embodied AI. In the near future, the massive haptic signals -encompassing both kinesthetic and vibrotactile signals- will place significant demands on both communication and computing resources. To address this challenge, we propose the first task-oriented sematic codec of low-delay vibrotactile transmission, namely, vibrotactile semantic codec (VTSC). Specifically, we design a perception-based vibrotactile semantic extraction mechanism (PSEM) that considers the high and low thresholds of vibrotactile perception in effective semantic coding while adhering to the low delay constraint. Inspired by this principle, we then propose a vibrotactile semantic encoder (VSE) with local and global semantic extractors, which can efficiently extract and preserve semantic features within the short frame context. Besides, we present a semantic distribution loss function to enhance the learning of meaningful representations. Comprehensive experiments demonstrate the superiority of our VTSC, achieving significantly higher task accuracy than the state-of-the-art vibrotactile codecs at the same compression ratio (CR), e.g. 60% improvement when CR=256. When compared to transferred audio-visual sematic codecs, our VTSC also shows promising improvements, validating the effectiveness our approach.

Junqi Liao, Yaojun Wu 0001, Chaoyi Lin, Zhipin Deng, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001

Neural video codecs (NVCs), leveraging the power of end-to-end learning, have demonstrated remarkable coding efficiency improvements over traditional video codecs. Recent research has begun to pay attention to the quality structures in NVCs, optimizing them by introducing explicit hierarchical designs. However, less attention has been paid to the reference structure design, which fundamentally should be aligned with the hierarchical quality structure. In addition, there is still significant room for further optimization of the hierarchical quality structure. To address these challenges in NVCs, we propose EHVC, an efficient hierarchical neural video codec featuring three key innovations: (1) a hierarchical multi-reference scheme that draws on traditional video codec design to align reference and quality structures, thereby addressing the reference-quality mismatch; (2) a lookahead strategy to utilize an encoder-side context from future frames to enhance the quality structure; (3) a layer-wise quality scale with random quality training strategy to stabilize quality structures during inference. With these improvements, EHVC achieves significantly superior performance to the state-of-the-art NVCs. Code will be released in: https://github.com/bytedance/NEVC.

Ahmad Alhilal, Ze Wu 0006, Teemu Kämäräinen, Tristan Braud, Matti Siekkinen

Virtual reality (VR) cloud gaming is increasingly developing in the gaming industry. Yet, the performance of the congestion control algorithms on top of which these systems build remains under-explored. In this study, we implement two industry-standard network congestion control algorithms, Google Congestion Control (GCC) and Network-Assisted Dynamic Adaptation (NADA), according to their Requests for Comments (RFCs), and integrate them into an open-source VR gaming system (ALVR). Including ALVR's congestion control (ALVR-ABR), we conduct extensive experiments on real-world networks to evaluate each algorithm's frame latency, target-to-receiving bitrate gap, dropped frames, image quality, and fairness among heterogeneous competing flows. GCC decreases frame latency by 35% compared to NADA and by 42% compared to ALVR. NADA and ALVR-ABR present significant gaps between the selected and received bitrate, causing substantial congestion-induced frame drops, while GCC has a minimal gap, resulting in minor frame drops, suggesting its suitability for game-player interaction. GCC exhibits a 2.7% and 5% decrease in image quality compared to NADA and ALVR-ABR, respectively, indicating slight immersion degradation. However, only NADA ensures a fair bandwidth share against loss-based flows due to its bitrate response to loss-induced congestion signals and lower sensitivity to delay gradients compared to GCC.

Jingrou Wu, Haoxian Liu, Jin Zhang 0001, Dan Wang 0002, Jing Jiang 0002

Volumetric videos are essential for immersive applications due to their engaging and realistic experiences. However, streaming them in real time over constrained, fluctuating networks remains challenging. Progressive streaming is an effective method to mitigate this issue by gradually enhancing video quality through incremental data transmission. However, existing progressive volumetric streaming solutions often rely on specific compression algorithms or require codec modifications, leading to poor compatibility with standard codecs. In this paper, we propose P2VS, a progressive partition-based volumetric video streaming framework, to achieve codec-independent progressive streaming. Specifically, P2VS leverages the unique structure of point cloud-based volumetric video to incrementally enhance video quality without being constrained by specific compression algorithms. Moreover, we propose adaptive streaming algorithms under this framework to enhance the quality of experience (QoE). Extensive simulations demonstrate that P2VS improves QoE by 21% on average compared to non-progressive streaming schemes. It also achieves better bandwidth efficiency and full compatibility with standard codecs. A prototype is built to verify the feasibility of P2VS.

Jiaxun Zhang, Haicheng Liao, Yumu Xie, Chengyue Wang 0001, Yanchen Guan, Bin Rao 0003, Zhenning Li 0001

Accurate accident anticipation remains challenging when driver cognition and dynamic road conditions are underrepresented in predictive models. In this paper, we propose CAMERA (Context-Aware Multi-modal Enhanced Risk Anticipation), a multi-modal framework integrating dashcam video, textual annotations, and driver attention maps for robust accident anticipation. Unlike existing methods that rely on static or environment-centric thresholds, CAMERA employs an adaptive mechanism guided by scene complexity and gaze entropy, reducing false alarms while maintaining high recall in dynamic, multi-agent traffic scenarios. A hierarchical fusion pipeline with Bi-GRU (Bidirectional GRU) captures spatio-temporal dependencies, while a Geo-Context Vision-Language module translates 3D spatial relationships into interpretable, human-centric alerts. Evaluations on the DADA-2000 and benchmarks show that CAMERA achieves state-of-the-art performance, improving accuracy and lead time. These results demonstrate the effectiveness of modeling driver attention, contextual description, and adaptive risk thresholds to enable more reliable accident anticipation.

Zhaohui Jiang, Xuening Feng, Tianchi Huang, Ruixiao Zhang, Paul Weng, Yifei Zhu 0001

Existing quality of experience (QoE)-driven adaptive bitrate (ABR) algorithms either fail to consider personalized QoE or rely on over-simplified QoE models, all resulting in unsatisfactory streaming experiences. Recognizing the wide existence of user feedback schemes in existing streaming applications, we introduce Q+, a framework leveraging progressively gathered personal user opinion scores from multiple interaction sessions for enhanced user-system alignment. Q+ first innovates QoE modeling by incorporating both pairwise ordinal and cardinal preferences constructed from scores. The capturing of both preferences ensures reliable and robust preference representation. Moreover, we design a monotonic neural network as the QoE model to capture the inherent monotonicity property in ABR services, improving model expressivity and generalization ability even with limited human feedback. To align the policy with the progressively updated QoE, we then develop a value-based reinforcement learning (RL) algorithm for bitrate control that integrates reward relabeling and calibrated prioritized experience replay. Extensive experiments reveal that Q+ consistently surpasses state-of-the-art rule-based, control-based, and RL-based baselines within only three sessions, improving QoE by 5.69% to 29.39% across diverse network conditions.

Lianchen Jia, Chaoyang Li 0002, Ziqi Yuan, Jiahui Chen 0009, Tianchi Huang, Jiangchuan Liu, Lifeng Sun

Over the past decade, adaptive video streaming technology has witnessed significant advancements, particularly driven by the rapid evolution of deep learning techniques. However, the black-box nature of deep learning algorithms presents challenges for developers in understanding decision-making processes and optimizing for specific application scenarios. Although existing research has enhanced algorithm interpretability through decision tree conversion, interpretability does not directly equate to developers' subjective comprehensibility. To address this challenge, we introduce ComTree, the first bitrate adaptation algorithm generation framework that considers comprehensibility. The framework initially generates the complete set of decision trees that meet performance requirements, then leverages large language models to evaluate these trees for developer comprehensibility, ultimately selecting solutions that best facilitate human understanding and enhancement. Experimental results demonstrate that ComTree significantly improves comprehensibility while maintaining competitive performance, showing potential for further advancement. The source code and appendix are available at https://github.com/thu-media/ComTree.

Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Hesheng Sun, Yu Liang 0001, Zhuzhong Qian, Han Zheng, Xiaokun Wang 0002, Ning Jiang

Video analytics pipelines migrating to edge deployments are facing performance bottlenecks under limited bandwidth. Non-uniform intra-frame encoding emerges to further compress pixels without affecting the output of the server deep neural network (DNN), while it is inefficient in high-resolution video streaming at low bandwidth. The detail enhancement capability of neural super-resolution (SR) permits resolution downsampling and aggressive compression on edge devices for low-latency transmission. To exploit its accuracy potential, DNN-oriented non-uniform encoding is expected to be additionally aware of SR models. However, traditional codecs struggle to cope with both quality optimization for SR and global semantic features for DNN. We advocate neural codecs for coordinated encoding and enhancement, enabling analytic-oriented video streaming with optimal accuracy-delay tradeoffs. Our system, VidIQ, achieves quality-enhanced real-time video analytics by 1) improving the network architecture of neural codecs (at two granularity) to integrate SR models into a DNN-oriented analytics pipeline, and 2) adapting the multi-scale encoder and SR-decoder to scene dynamics (i.e., content and bandwidth variations) with the help of the monolithic controller to hold a performance advantage. Extensive evaluations showcase that VidIQ reduces end-to-end delay by 35.8% and improves analytics accuracy by 21.2% compared to the recent video compression, enhancement, and streaming baselines.