论文检索

输入标题、作者或关键词,从 8,216 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
8,216篇论文匹配“New Approaches”
第 103 / 411 页

Jiahua Dong, Hui Yin, Wenqi Liang, Hanbin Zhao, Henghui Ding, Nicu Sebe, Salman Khan, Fahad Shahbaz Khan

Video instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume that the categories of object instances remain fixed over time. Moreover, they experience catastrophic forgetting of old classes when required to continuously learn object instances belonging to new categories. To resolve these challenges, we develop a novel Hierarchical Visual Prompt Learning (HVPL) model that overcomes catastrophic forgetting of previous categories from both frame-level and video-level perspectives. Specifically, to mitigate forgetting at the frame level, we devise a task-specific frame prompt and an orthogonal gradient correction (OGC) module. The OGC module helps the frame prompt encode task-specific global instance information for new classes in each individual frame by projecting its gradients onto the orthogonal feature space of old classes. Furthermore, to address forgetting at the video level, we design a task-specific video prompt and a video context decoder. This decoder first embeds structural inter-class relationships across frames into the frame prompt features, and then propagates task-specific global video contexts from the frame prompt features to the video prompt. Through rigorous comparisons, our HVPL model proves to be more effective than baseline approaches.

Fatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom O. Ikezogwo, Beibin Li, Tejoram Vivekanandan, Joann G. Elmore, Ranjay Krishna, Linda Shapiro

Diagnosing diseases through histopathology whole slide images (WSIs) is fundamental in modern pathology but is challenged by the gigapixel scale and complexity of WSIs. Trained histopathologists overcome this challenge by navigating the WSI, looking for relevant patches, taking notes, and compiling them to produce a final holistic diagnostic. Traditional AI approaches, such as multiple instance learning and transformer-based models, fail short of such a holistic, iterative, multi-scale diagnostic procedure, limiting their adoption in the real-world. We introduce PathFinder, a multi-modal, multi-agent framework that emulates the decision-making process of expert pathologists. PathFinder integrates four AI agents, the Triage Agent, Navigation Agent, Description Agent, and Diagnosis Agent, that collaboratively navigate WSIs, gather evidence, and provide comprehensive diagnoses with natural language explanations. The Triage Agent classifies the WSI as benign or risky; if risky, the Navigation and Description Agents iteratively focus on significant regions, generating importance maps and descriptive insights of sampled patches. Finally, the Diagnosis Agent synthesizes the findings to determine the patient's diagnostic classification. Our Experiments show that PathFinder outperforms state-of-the-art methods in skin melanoma diagnosis by 8% while offering inherent explainability through natural language descriptions of diagnostically relevant patches. Qualitative analysis by pathologists shows that the Description Agent's outputs are of high quality and comparable to GPT-4o. PathFinder is also the first AI-based system to surpass the average performance of pathologists in this challenging melanoma classification task by 9%, setting a new record for efficient, accurate, and interpretable AI-assisted diagnostics in pathology. Our Data, code and models will be made available.

Maximilian Pittner, Joel Janai, Mario Faigle, Alexandru Paul Condurache

3D lane detection has emerged as a critical challenge in autonomous driving, encompassing identification and localization of lane markings and the 3D road surface. Conventional 3D methods detect lanes from dense birds-eye-viewed (BEV) features, though erroneous transformations often result in a poor feature representation misaligned with the true 3D road surface. While recent sparse lane detectors have surpassed dense BEV approaches, they completely disregard valuable lane-specific priors. Furthermore, existing methods fail to utilize historic lane observations, which yield the potential to resolve ambiguities in situations of poor visibility. To address these challenges, we present SparseLaneSTP, a novel method that integrates both geometric properties of the lane structure and temporal information into a sparse lane transformer. It introduces a new lane-specific spatio-temporal attention mechanism, a continuous lane representation tailored for sparse architectures as well as temporal regularization. Identifying weaknesses of existing 3D lane datasets, we also introduce a precise and consistent 3D lane dataset using a simple yet effective auto-labeling strategy. Our experimental section proves the benefits of our contributions and demonstrates state-of-the-art performance across all detection and error metrics on existing 3D lane detection benchmarks as well as on our novel dataset.

Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Kristen Grauman

We introduce Switch-a-View, a model that learns to automatically select the viewpoint to display at each timepoint when creating a how-to video. The key insight of our approach is how to train such a model from unlabeled--but human-edited--video samples. We pose a pretext task that pseudo-labels segments in the training videos for their primary viewpoint (egocentric or exocentric), and then discovers the patterns between the visual and spoken content in a how-to video on the one hand and its view-switch moments on the other hand. Armed with this predictor, our model can be applied to new multi-view videos to orchestrate which viewpoint should be displayed when. We demonstrate our idea on a variety of real-world videos from HowTo100M and Ego-Exo4D, and rigorously validate its advantages.

Jinghan You, Shanglin Li, Yuanrui Sun, Jiangchuan Wei, Mingyu Guo, Chao Feng, Jiao Ran

Vision Transformers (ViTs) have revolutionized large-scale visual modeling, yet remain underexplored in face recognition (FR) where CNNs still dominate. We identify a critical bottleneck: CNN-inspired training paradigms fail to unlock ViT's potential, leading to suboptimal performance and convergence instability.To address this challenge, we propose LVFace, a ViT-based FR model that integrates Progressive Cluster Optimization (PCO) to achieve superior results. Specifically, PCO sequentially applies negative class sub-sampling (NCS) for robust and fast feature alignment from random initialization, feature expectation penalties for centroid stabilization, performing cluster boundary refinement through full-batch training without NCS constraints. LVFace establishes a new state-of-the-art face recognition baseline, surpassing leading approaches such as UniFace and TopoFR across multiple benchmarks. Extensive experiments demonstrate that LVFace delivers consistent performance gains, while exhibiting scalability to large-scale datasets and compatibility with mainstream VLMs and LLMs. Notably, LVFace secured 1st place in the ICCV 2021 Masked Face Recognition (MFR)-Ongoing Challenge (March 2025), proving its efficacy in real-world scenarios. Project is available at https://github.com/bytedance/LVFace.

Haochen Wang, Yucheng Zhao, Tiancai Wang, Haoqiang Fan, Xiangyu Zhang, Zhaoxiang Zhang

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant obstacle. To address this issue, typical approaches focus on injecting 3D awareness into 2D LMMs by designing 3D input-level scene representations. This work provides a new perspective. We introduce reconstructive visual instruction tuning with 3D-awareness (ROSS3D), which integrates 3D-aware visual supervision into the training procedure. Specifically, it incorporates cross-view and global-view reconstruction. The former requires reconstructing masked views by aggregating overlapping information from other views. The latter aims to aggregate information from all available views to recover Bird's-Eye-View images, contributing to a comprehensive overview of the entire scene. Empirically, ROSS3D achieves state-of-the-art performance across various 3D scene understanding benchmarks. More importantly, our semi-supervised experiments demonstrate significant potential in leveraging large amounts of unlabeled 3D vision-only data.

Kaijie Yin, Zhiyuan Zhang, Shu Kong, Tian Gao, Cheng-Zhong Xu, Hui Kong

In this paper, we propose Binarized Change Detection (BiCD), the first binary neural network (BNN) designed specifically for change detection. Conventional network binarization approaches, which directly quantize both weights and activations in change detection models, severely limit the network's ability to represent input data and distinguish between changed and unchanged regions. This results in significantly lower detection accuracy compared to real-valued networks. To overcome these challenges, BiCD enhances both the representational power and feature separability of BNNs, improving detection performance. Specifically, we introduce an auxiliary objective based on the Information Bottleneck (IB) principle, guiding the encoder to retain essential input information while promoting better feature discrimination. Since directly computing mutual information under the IB principle is intractable, we design a compact, learnable auxiliary module as an approximation target, leading to a simple yet effective optimization strategy that minimizes both reconstruction loss and standard change detection loss.Extensive experiments on street-view and remote sensing datasets demonstrate that BiCD establishes a new benchmark for BNN-based change detection, achieving state-of-the-art performance in this domain.

Junsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu, Sayak Paul, Junyu Chen, Han Cai, Song Han, Enze Xie

This paper presents SANA-Sprint, an efficient diffusion model for ultra-fast text-to-image (T2I) generation. SANA-Sprint is built on a pre-trained foundation model and augmented with hybrid distillation, dramatically reducing inference steps from 20 to 1-4.We introduce three key innovations: (1) We propose a training-free approach that transforms a pre-trained flow-matching model for continuous-time consistency distillation (sCM), eliminating costly training from scratch and achieving high training efficiency. Our hybrid distillation strategy combines sCM with latent adversarial distillation (LADD): sCM ensures alignment with the teacher model, while LADD enhances single-step generation fidelity. (2) SANA-Sprint is a unified step-adaptive model that achieves high-quality generation in 1-4 steps, eliminating step-specific training and improving efficiency. (3) We integrate ControlNet with SANA-Sprint for real-time interactive image generation, enabling instant visual feedback for user interaction. SANA-Sprint establishes a new Pareto frontier in speed-quality tradeoffs, achieving state-of-the-art performance with 7.59 FID and 0.74 GenEval in just 1 step -- outperforming FLUX-schnell (7.94 FID / 0.71 GenEval) while being 10x faster (0.1s vs 1.1s on H100). It also achieves 0.1s (T2I) and 0.25s (ControlNet) latency for 1024x1024 images on H100, and 0.31s (T2I) on an RTX 4090, showcasing its exceptional efficiency and potential for AI-powered consumer applications (AIPC). Code and pre-trained models will be open-sourced.

Kai Tong, Kang Pan, Xiao Zhang, Erli Meng, Run He, Yawen Cui, Nuoyan Guo, Huiping Zhuang

Large Language Models (LLMs) possess encompassing capabilities that can process diverse language-related tasks. However, finetuning on LLMs will diminish this general skills and continual finetuning will further cause severe degradation on accumulated knowledge. Recently, Continual Learning (CL) in Large Language Models (LLMs) arises which aims to continually adapt the LLMs to new tasks while maintaining previously learned knowledge and inheriting general skills. Existing techniques either leverage previous data to replay, leading to extra computational costs, or utilize a single parameter-efficient module to learn the downstream task, constraining new knowledge absorption with interference between different tasks. Toward these issues, this paper proposes Analytic Subspace Routing(Any-SSR) to address these challenges. For each task, we isolate the learning within a subspace of deep layers' features via low-rank adaptation, eliminating knowledge interference between different tasks. Additionally, we propose an analytic routing mechanism to properly utilize knowledge learned in different subspaces. Our approach employs Recursive Least Squares to train a multi-task router model, allowing the router to dynamically adapt to incoming data without requiring access to historical data. Also, the router effectively assigns the current task to an appropriate subspace and has a non-forgetting property of previously learned tasks with a solid theoretical guarantee. Experimental results demonstrate that our method achieves near-perfect retention of prior knowledge while seamlessly integrating new information, effectively overcoming the core limitations of existing methods. Our code is available at https://github.com/ZHUANGHP/Any-SSR

Tomasz Niewiadomski, Anastasios Yiannakidis, Hanz Cuevas-Velasquez, Soubhik Sanyal, Michael J. Black, Silvia Zuffi, Peter Kulits

The model-based estimation of 3D animal pose and shape from images enables computational modeling of animal behavior. Training models for this purpose requires large amounts of labeled image data with precise pose and shape annotations. However, capturing such data requires the use of multi-view or marker-based motion-capture systems, which are impractical to adapt to wild animals in situ and impossible to scale across a comprehensive set of animal species. Some have attempted to address the challenge of procuring training data by pseudo-labeling individual real-world images through manual 2D annotation, followed by 3D-parameter optimization to those labels. While this approach may produce silhouette-aligned samples, the obtained pose and shape parameters are often implausible due to the ill-posed nature of the monocular fitting problem. Sidestepping real-world ambiguity, others have designed complex synthetic-data-generation pipelines leveraging game engines and collections of artist-designed 3D assets. Such engines yield perfect ground-truth labels but are often lacking in visual realism and require considerable manual effort to adapt to new species or environments. We propose an alternative approach to synthetic-data generation: rendering with a conditional image-generation model. We introduce a pipeline that samples a diverse set of poses and shapes for a variety of mammalian quadrupeds and generates realistic images with corresponding ground-truth pose and shape parameters. To demonstrate the scalability of our approach, we introduce GenZoo, a synthetic dataset containing one million images of distinct subjects. We train a 3D pose and shape regressor on GenZoo, which achieves state-of-the-art performance on a real-world multi-species 3D animal pose and shape estimation benchmark, despite being trained solely on synthetic data. We release our data and pipeline at https://genzoo.is.tue.mpg.de.

Zhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue, Xinchao Wang

We present OminiControl, a novel approach that rethinks how image conditions are integrated into Diffusion Transformer (DiT) architectures. Current image conditioning methods either introduce substantial parameter overhead or handle only specific control tasks effectively, limiting their practical versatility. OminiControl addresses these limitations through three key innovations: (1) a minimal architectural design that leverages the DiT's own VAE encoder and transformer blocks, requiring just 0.1% additional parameters; (2) a unified sequence processing strategy that combines condition tokens with image tokens for flexible token interactions; and (3) a dynamic position encoding mechanism that adapts to both spatially-aligned and non-aligned control tasks. Our extensive experiments show that this streamlined approach not only matches but surpasses the performance of specialized methods across multiple conditioning tasks. To overcome data limitations in subject-driven generation, we also introduce Subjects200K, a large-scale dataset of identity-consistent image pairs synthesized using DiT models themselves. This work demonstrates that effective image control can be achieved without architectural complexity, opening new possibilities for efficient and versatile image generation systems.

Minh Tran, Hongda Mao, Qingshuang Chen, Yelin Kim

Generating body pose from head-mounted, egocentric inputs is essential for immersive VR/AR and assistive technologies, as it supports more natural interactions. However, the task is challenging due to limited visibility of body parts in first-person views and the sparseness of sensory data, with only a single device placed on the head. To address these challenges, we introduce Head2Body, a novel framework for body pose estimation that effectively combines IMU and visual data. First, we introduce a pre-trained IMU encoder, trained on over 1,700 hours of head-IMU data from wearable eyeglasses, to better capture detailed temporal motion cues given limited labeled egocentric pose data. For visual processing, we leverage large vision-language models (LVLMs) to segment body parts that appear sporadically in video frames to improve visual feature extraction. To better guide the pose generation process with sparse signals from only head-mounted devices, we incorporates a Vector Quantized Variational Autoencoder (VQ-VAE) to represent poses as discrete tokens, which capture high-frequency motion patterns and provide a more structured representation of body pose. Our experiments demonstrate the effectiveness of the proposed approach, yielding 6-13% gains over state-of-the-art baselines on four datasets: AMASS, KinPoly, GIMO, and EgoExo4D. By capturing subtle temporal dynamics and leveraging complementary sensory data, our approach advances accurate egocentric body pose estimation and sets a new benchmark for multi-modal, first-person motion tracking.

Ge Zheng, Jiaye Qian, Jiajin Tang, Sibei Yang

Large Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-form responses, often attributed to accumulated uncertainties. In this paper, we ask: Does increased hallucination result solely from length-induced errors, or is there a deeper underlying mechanism? After a series of preliminary experiments and findings, we suggest that the risk of hallucinations is not caused by length itself but by the increased reliance on context for coherence and completeness in longer responses. Building on these insights, we propose a novel "induce-detect-suppress" framework that actively induces hallucinations through deliberately designed contexts, leverages induced instances for early detection of high-risk cases, and ultimately suppresses potential object-level hallucinations during actual decoding. Our approach achieves consistent, significant improvements across all benchmarks, demonstrating its efficacy. The strong detection and improved hallucination mitigation not only validate our framework but, more importantly, re-validate our hypothesis on context. Rather than solely pursuing performance gains, this study aims to provide new insights and serves as a first step toward a deeper exploration of hallucinations in LVLMs' longer responses.

Wenlong Luo, Shizhou Zhang, De Cheng, Yinghui Xing, Guoqiang Liang, Peng Wang, Yanning Zhang

Incremental object detection (IOD) is crucial for enabling AI systems to continuously learn new object classes over time while retaining knowledge of previously learned categories, allowing model to adapt to dynamic environments without forgetting prior information.Existing IOD methods primarily employ knowledge distillation to mitigate catastrophic forgetting, yet these approaches overlook class overlap issues, often resulting in suboptimal performance. In this paper, we propose a novel framework for IOD that leverages a decoupled gradient alignment technique on top of the specially proposed pseudo-labeling strategy. Our method employs a Gaussian Mixture Model to accurately estimate pseudo-labels of previously learned objects in current training images, effectively functioning as a knowledge-replay mechanism. This strategy reinforces prior knowledge retention and prevents the misclassification of unannotated foreground objects from earlier classes as background. Furthermore, we introduce an adaptive gradient decomposition and alignment method to maintain model stability while facilitating positive knowledge transfer. By aligning gradients from both old and new classes, our approach preserves previously learned knowledge while enhancing plasticity for new tasks. Extensive experiments on two IOD benchmarks demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods.

Balu Adsumilli, Jianle Chen, In Suk Chong, Yilin Wang 0001

Video processing and compression are being re-envisioned in the age of AI. Traditional video codecs, which rely on rigid, pre-defined rules, are being augmented and, in some cases, replaced by AI-driven approaches. These new methods leverage machine learning to intelligently analyze video content, allowing for more adaptive and efficient compression. We are going to discuss AOM's new codec AV2, its low- and high-level features that enable significantly smaller file sizes with no perceptible loss in quality, a crucial development for streaming and storage. The shift to AI has also transformed how we evaluate video quality. Traditional metrics, while directionally useful, don't always align with human perception, especially for user generated content (UGC). They fail to capture what's most important for machine vision tasks. We will talk about new AI-based quality metrics that are being developed. They correlate better with a human's subjective experience and a machine's ability to perform tasks like object recognition. Along the way, we'll cover large scale industrial infrastructure challenges and the ways to achieve high reliability and accuracy.

Yasunori Mochizuki

NEC is the leading ICT technology provider in the B-to-B market and is actively integrating cutting-edge technologies into its business solutions to drive innovation, enhance capabilities, and create new value for its customers in a broad spectrum of industrial segments. And the recent business focus of NEC is to support digital transformation of business processes of customer enterprises by leveraging technical capabilities in AI, Cyber Security and Communication. This keynote discusses the specific role of NEC's Research in such a business context by sharing a variety of generative and multimodal AI-related use cases that are aimed at solving critical customer challenges in the real-world. From the multimedia perspective, the topics will include world-leading facial recognition technology for security boost and enhanced customer experience, development of drive-recorder video analytics for insurance adjusters leveraging visual language model (VLM) and medical document generation AI service for genuinely supporting overworked clinical doctors. Meanwhile, distributed acoustic sensing technology using optical fiber cables is opening a new opportunity for infrastructure and incident monitoring solutions after integration with AI and ML algorithms. As the common denominator, our commitment of solving critical customer challenges requires (and justifies) nurturing both world-class excellence in performing academic research and accumulated experience and/or culture of application-oriented technology refinement as well as technology combination to ensure business-ready practicality. Also, being the industrial research organization, we are engaged at the forefront of customer co-creation and co-design that play an indispensable role in pinpointing customer's critical challenges. These expertise and practices are indeed the core ingredients of NEC's Research for creating new business opportunities from the technology innovation approach. Furthermore, we also envision that such an industrial lab model in the Generative AI era will become the driver of a new technology paradigm - industry segment-oriented customizable foundation models and business transforming Agentic AI framework.

Aleksandr Farseev

Recent industry discussions, led by Meta CEO Mark Zuckerberg, envision a radical future where advertisers provide only goals and budgets, while AI systems autonomously manage the entire advertising pipeline - creative, targeting, optimization, and measurement. Zuckerberg described this as a ''redefinition of the category of advertising,'' promising speed, scalability, and simplicity. However, critics warn of risks to transparency and trust, as such systems could displace agencies and in-house teams, leaving brands with little oversight of the ''black box'' outcomes. Alternative approaches frame AI in marketing as a collaborator. Google's ACAI, built with DeepMind and Oxford, treats AI as a co-creator: advertisers provide assets and insights, which are synthesized into structured ''super-prompts'' to guide campaigns, ensuring agency and transparency. Meta's AI Sandbox takes a similar path, offering generative tools for copy, visuals, and resizing while preserving human input. Evidence shows such co-created campaigns deliver higher returns on ad spend, highlighting the value of AI-human collaboration. A third path, seen in platforms like SOMIN, Nielsen, and GWI, emphasizes insight and explainability over automation. Rather than generating ads, these systems analyze successful campaigns to reveal audience, creative, and optimization patterns, embedding AI as a data-driven guide to human creativity. Together, these models illuminate three diverging philosophies: (i) replacement, where AI supplants creative and strategic functions (Meta); (ii) empowerment, where AI scaffolds human creativity (Google); and (iii) insight, where AI explains and enhances strategic thinking (SOMIN and peers). AI adoption is reshaping agencies: while tasks like resizing, reporting, and A/B testing are automated, their real value lies in guiding AI to ensure brand alignment, cultural nuance, and authenticity. This shift demands hybrid skills that merge storytelling with data fluency, AI literacy, and prompt engineering. To adapt, agencies are investing in large-scale upskilling - WPP alone delivered over 150,000 AI training sessions in 2025 - and new roles such as creative technologists, AI trainers, and ethics officers are emerging, positioning agencies as curators of AI-driven creativity rather than passive executors. Hybrid agencies combining storytelling, data science, and engineering - like Accenture Song with its multi-billion AI investments - are disrupting traditional models by offering end-to-end capabilities once spread across multiple firms. Analysts call them ''talent magnets,'' attracting professionals eager to work across disciplines. Meanwhile, Meta's automation push threatens traditional roles, forcing agencies to compete through culture, talent, and authenticity, positioning themselves as AI facilitators that keep human creativity central.

Qiang Sheng 0001, Peng Qi 0005, Tianyun Yang, Yuyan Bu, Wynne Hsu, Mong-Li Lee, Juan Cao 0001

Recent progress of generative AI and the popularity of short-form video-sharing platforms have raised new risks of misinformation video issues, posing a potential threat to online multimedia ecosystems. With the aid of generative AI tools, producing and spreading vivid, persuasive misinformation videos has been easier, while detecting and preventing them has become harder. This tutorial introduces how to characterize, detect, and prevent misinformation videos, which consists of three technical parts: 1) Characterization of AI-generated and human-edited misinformation videos; 2) Detection approaches, covering those tailored for fully generated, manipulated, and human-edited videos; and 3) Prevention strategies, including those effective for the creation and spread phases. This tutorial concludes by discussing the status quo and ongoing challenges and highlighting the promising directions for future research. We expect to bring broader attention to misinformation video issues, gather and communicate with researchers of interest, and facilitate the engagement of those who are new to this field.

Cheng Jin 0001, Mingli Song, Rui Wang 0032, Xingjiao Wu

This workshop addresses next-generation methods in multimedia research, with a focus on content generation, quality assessment, and dataset development. These three areas are foundational for advancing multimedia technologies and applications. Emerging approaches in multimedia content generation, powered by generative AI and multimodal learning, are reshaping domains such as entertainment, advertising, education, and healthcare. At the same time, robust quality assessment is essential to ensure that generated content achieves high standards of perceptual fidelity, semantic consistency, and user satisfaction, thereby determining the real-world impact of multimedia systems. Datasets remain indispensable for training and evaluating algorithms, and innovative strategies in dataset construction-ranging from augmentation and annotation to addressing issues of bias and small-sample imbalance-are driving the development of more reliable and ethical multimedia applications. By convening leading researchers and practitioners, this workshop provides a platform to explore state-of-the-art methods, share best practices, and discuss open challenges in next-generation multimedia research. The goal is to foster interdisciplinary collaboration and inspire innovative solutions that advance the creation, evaluation, and application of multimedia content, setting new benchmarks for the field and shaping the future of multimedia technologies.

Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001

The 8th ACM International Workshop on Multimedia Content Analysis in Sports is held in Dublin, Ireland on October 28th, 2025. It is co-located with ACM Multimedia 2025. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing the multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation as well as understanding, statistical analysis, and evaluation in amateur and professional sports. There is a lack of research communities focusing on the fusion of multiple modalities. Thus, this workshop series on multimedia content analysis in sports aims to contribute to the closure of this research gap by bringing together the breadth and depth of these diverse approaches to stimulate each other with new ideas and foster research progress.