论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 3 / 81 页

Tai Tan Mai, Allie Tran, Quang-Linh Tran, An Nguyen, Hoang Nguyen, Tho Quan, Duc-Tien Dang-Nguyen, Cathal Gurrin

The Second ACM Workshop on AI-Powered Question & Answering Systems for Multimedia (AIQAM'25) was held on 27 October 2025 in Dublin, Ireland, co-located with ACM Multimedia 2025. The workshop's main objective is to create a collaborative and inclusive space for researchers at the intersection of Artificial Inteligence (AI), large language models (LLMs), multimodal information retrieval, and question answering systems. Building on the success of its first edition (AIQAM'24) at ICMR 2024, AIQAM'25 provided a forum for presenting novel methods, applications, and surveys, which address the challenges of integrating text, image, audio, and video data into QA systems. The programme featured contributions ranging from methodological advances in reasoning and evaluation frameworks, domain-specific applications in education and sustainability, to a survey work reviewing the state of multimedia retrieval-augmented QA. This summary paper outlines the objectives and scope of the workshop, describes its format, highlights the keynote and accepted contributions, and acknowledges the efforts of the organising committee.

Wei Jiang 0001, Zhenghao Chen, Dong Xu 0001

The goal of this workshop is to showcase the latest advancements in generative AI (GAI) for creating, editing, restoring, and compressing rich media data, including images, videos, and 3D content. GAI models such as VAEs, GANs, and diffusion models have demonstrated remarkable impact in both academic research and industrial applications. For example, GAI enables users to design and generate synthetic yet realistic content without requiring professional artistic or technical expertise, driving significant market growth in gaming and entertainment. Beyond creative applications, GAI also provides crucial simulated data for training embodied AI agents. When applied to media restoration and synthesis, GAI techniques can further alleviate transmission challenges by offloading computation to client devices. To advance this field, the workshop will host four competition tracks using novel industry-level data, solicit high-quality paper submissions, and invite leading speakers from academia and industry to foster collaboration and innovation. In particular, the competition focuses on media generation and transmission with GAI. The first three tracks address reducing computation and transmission costs for efficient media delivery, while the fourth track focuses on controlled novel content creation. To support these challenges, a large-scale multi-modality, multi-view dataset named M3VIR is provided. This dataset comprises a diverse collection of videos simulated using the UE5 Unreal Engine, with carefully matched content serving as ground truth for the competition tasks.

Wei Zhou 0021, Hadi Amirpour, Li Yu 0004, Jungong Han, Richang Hong, Paul L. Rosin

Recent years have witnessed an unprecedented growth of multimodal data in healthcare, ranging from distributed sensors and medical imaging devices (MRI, CT, X-rays) to digital health platforms that integrate audio, video, 3D geometry, and clinical text. The increasing availability of such data presents significant opportunities for computer-aided diagnosis and intelligent healthcare solutions, yet also poses substantial challenges in multimodal integration, large-scale analysis, and real-world deployment. The 2nd International Workshop on Multimedia Computing for Health and Medicine (MCHM'25), held in conjunction with ACM Multimedia 2025, focuses on advanced multimedia computing techniques, including mobile and hardware solutions, for tackling real-world problems in healthcare. The workshop brings together researchers and practitioners in multimedia computing, artificial intelligence, and medicine to explore emerging methods, applications, and systems that have a direct impact on human health.

Aik Beng Ng, Yethoven Tukimin, Jeannie S. Lee, Megani Rajendran, Chek Tien Tan, Indriyati Atmosukarto

This workshop is part of the ACM Multimedia 2025 Conference and is organized by the ACM I2M Chapter, consisting of both industry and academia members. The rapid convergence of Artificial Intelligence (AI), Human-Computer Interaction (HCI), and immersive multimedia is redefining the landscape of intelligent and adaptive digital experiences. As ACM Multimedia 2025 emphasizes cutting-edge multimedia systems, this workshop directly contributes to its vision by exploring AI's transformative role in immersive media. Through AI-driven multimedia interaction, adaptive virtual environments, and intelligent content generation, this workshop will showcase how AI is enhancing the creation and experience of digital worlds. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3728487

Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001

The 8th ACM International Workshop on Multimedia Content Analysis in Sports is held in Dublin, Ireland on October 28th, 2025. It is co-located with ACM Multimedia 2025. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing the multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation as well as understanding, statistical analysis, and evaluation in amateur and professional sports. There is a lack of research communities focusing on the fusion of multiple modalities. Thus, this workshop series on multimedia content analysis in sports aims to contribute to the closure of this research gap by bringing together the breadth and depth of these diverse approaches to stimulate each other with new ideas and foster research progress.

Valérie Gouet-Brunet, Edgar Roman-Rangel, Li Weng

SUMAC 2025 is the 7th edition of the workshop on analySis, Understanding and proMotion of heritAge Contents. It is held in Dublin, Ireland, on 27 October and is co-located with the 33rd ACM International Conference on Multimedia. The workshop's objective is to present and discuss the latest and most significant trends, challenges, and advances in the fields of machine learning, signal processing, multimodal techniques, and human-machine interaction. The workshop is dedicated to the valorization of cultural heritage, with an emphasis on unlocking and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.

Quang-Linh Tran, Hoang-Bao Le, Thang-Long Nguyen-Ho, Graham Healy, Liting Zhou, Allie Tran

We present the DCU team's system for the CASTLE Challenge at ACM Multimedia 2025, which explores video retrieval and question answering in egocentric, multi-user environments. Our system adapts techniques developed for lifelogging, particularly event-based semantic retrieval and QA pipelines, to the CASTLE dataset with minimal architectural changes. It combines vision-language embeddings, transcript-based retrieval, and person tracking to support both automatic and interactive search workflows. In the interactive track, we introduce a modular interface for narrative reconstruction and exploratory search. Qualitative results show that the system can generate plausible, evidence-based answers to complex multimodal queries. These findings suggest that lifelog retrieval systems offer a viable foundation for broader egocentric video analysis.

Omar Shahbaz Khan, Ujjwal Sharma 0001, Gonçalo Marcelino, Aaron Duane, Stevan Rudinac, Marcel Worring, Björn Þór Jónsson 0001

We introduce multiXview, an interactive retrieval framework for synchronized multi-camera video collections. It features a multi-index search engine that supports natural-language queries over visual embeddings, speech transcripts, and scene descriptions. It supports a synchronized multi-stream player offering parallel playback, and a timeline-based navigation view for temporal scoping and faceted exploration. These components address the redundancy and fragmentation of overlapping egocentric and exocentric video feeds and enable users to locate, aggregate, and reconstruct events across partial perspectives. This paper focuses on system design and implementation, with quantitative and qualitative evaluation to take place at the CASTLE 2025 Grand Challenge Interactive Track.

Luca Rossetto, Werner Bailer, Cathal Gurrin, Duc-Tien Dang-Nguyen, Klaus Schoeffmann, Allie Tran

The inaugural edition of the CASTLE grand challenge was held at ACM Multimedia 2025. The focus of the CASTLE challenge is to advance the state-of-the-art in analysis and understanding of multimodal data, especially centered around multistream ego- and exo-centric video. In this first instance of the challenge, participants had to solve three types of tasks - event instance search, object instance search, and question answering - in a fully automatic or interactive setting.

Thinh-Phuc Nguyen, Thanh-Hai Nguyen, Gia-Huy Dinh, Lam-Huy Nguyen, Minh-Triet Tran, Trung-Nghia Le

Image captioning systems often produce generic descriptions that fail to capture event-level semantics which are crucial for applications like news reporting and digital archiving. We present ReCap, a novel pipeline for event-enriched image retrieval and captioning that incorporates broader contextual information from relevant articles to generate narrative-rich, factually grounded captions. Our approach addresses the limitations of standard vision-language models that typically focus on visible content while missing temporal, social, and historical contexts. ReCap comprises three integrated components: (1) a robust two-stage article retrieval system using DINOv2 embeddings with global feature similarity for initial candidate selection followed by patch-level mutual nearest neighbor similarity re-ranking; (2) a context extraction framework that synthesizes information from article summaries, generic captions, and original source metadata; and (3) a large language model-based caption generation system with Semantic Gaussian Normalization to enhance fluency and relevance. Evaluated on the OpenEvents V1 dataset as part of Track 1 in the EVENTA 2025 Grand Challenge, ReCap achieved a strong overall score of 0.54666, ranking 2nd on the private test set. These results highlight ReCap's effectiveness in bridging visual perception with real-world knowledge, offering a practical solution for context-aware image understanding in high-stakes domains. The code is available at https://github.com/Noridom1/EVENTA2025-Event-Enriched-Image-Captioning.

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

Event-based image retrieval from free-form captions presents a significant challenge: models must understand not only visual features but also latent event semantics, context, and real-world knowledge. Conventional vision-language retrieval approaches often fall short when captions describe abstract events, implicit causality, temporal context, or contain long, complex narratives. To tackle these issues, we introduce a multi-stage retrieval framework combining dense article retrieval, event-aware language model reranking, and efficient image collection, followed by caption-guided semantic matching and rank-aware selection. We leverage Qwen3 for article search, Qwen3-Reranker for contextual alignment, and Qwen2-VL for precise image scoring. To further enhance performance and robustness, we fuse outputs from multiple configurations using Reciprocal Rank Fusion (RRF). Our system achieves the top-1 score on the private test set of Track 2 in the EVENTA 2025 Grand Challenge, demonstrating the effectiveness of combining language-based reasoning and multimodal retrieval for complex, real-world image understanding. The code is available at https://github.com/vdkhoi20/EVENT-Retriever.

Nam-Quan Nguyen, Minh-Hoang Le 0001, Vinh-Toan Vong, Minh-Triet Tran

In many real-world applications, labeling an image ''a man riding a horse'' fails to satisfy demands for the who, when, where, and why. Although LVLMs excel at describing visual content, isolated images often lack the event context; users thus rely on related news articles or social posts to enrich them, but cropping or resizing complicates tracking back to their source. In this paper, we propose ENRIC, an innovative end-to-end system for the EVENTA Challenge Track 1, leveraging the OpenEvents-V1 dataset, comprising over 200,000 news articles paired with more than 400,000 images. Our system includes three components: (1) semantic retrieval filters candidate article images via vision-language embeddings, (2) uncertainty-guided re-ranking flags ambiguous queries using three confidence heuristics and re-ranks candidates by combining visual similarity with texture similarity, and (3) event-aware caption generation employs chain-of-thought prompting that aggregates five inputs from article, image, and CIDEr-derived contexts to guide the LLM in incorporating all necessary elements. ENRIC achieved the highest combined evaluation score of 0.5501, ranking first and outperforming other solutions across nearly all metrics. By combining semantic retrieval, uncertainty-guided re-ranking, and event-aware caption generation, ENRIC demonstrates the efficiency of its approach for event-enriched image analysis. GitHub repository: https://github.com/NamQuanProject/EVENTA25-ENRIC

Thien-Phuc Tran, Minh-Quang Nguyen, Minh-Triet Tran, Tam V. Nguyen 0002, Trong-Le Do, Duy-Nam Ly, Viet-Tham Huynh, Khanh-Duy Le, Mai-Khiem Tran, Trung-Nghia Le

The Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks largely focus on surface-level recognition of people, objects, and scenes, often overlooking the contextual and semantic dimensions that define real-world events. EVENTA addresses this gap by integrating contextual, temporal, and semantic information to capture the who, when, where, what, and why behind an image. Built upon the OpenEvents V1 dataset, the challenge features two tracks: Event-Enriched Image Retrieval and Captioning, and Event-Based Image Retrieval. A total of 45 teams from six countries participated, with evaluation conducted through Public and Private Test phases to ensure fairness and reproducibility. The top three teams were invited to present their solutions at ACM Multimedia 2025. EVENTA establishes a foundation for context-aware, narrative-driven multimedia AI, with applications in journalism, media analysis, cultural archiving, and accessibility. Further details about the challenge are available at the official homepage: https://ltnghia.github.io/eventa/eventa-2025.

Zhichao Xia, Yichi Zhang, Yanjun Chi, Lingsi Zhu, Mohan Jing, Jun Yu 0001

Micro-action refers to subtle, low-intensity non-verbal behaviors that can provide insights into an individual's underlying emotions and intentions. Due to its brief duration and significant overlap, identifying these micro-actions poses a challenge for current models. In response to these challenges, this paper proposes a novel multi-feature fusion framework, which extracts coarse-grained body features and fine-grained action features separately. Specifically, we present Temporal Contextualization for fine-grained learning, a cross-frame injection mechanism designed to capture essential spatio-temporal information and introduce a 3D-ResNet Adapter for coarse-grained learning, which aggregates temporal data and facilitates parameter-efficient fine-tuning. In consideration of the task dataset distribution's long-tail nature, the implementation of Feature Decoupling is undertaken, adopting a two-stage training strategy. By conducting experiments, the aforementioned hierarchical multi-feature extraction and aggregation approach has been demonstrated to yield substantial enhancement in Micro-Action Recognition. Our method attains an F1-mean score of 77.75% on the MA-52 dataset, ranking 1st in the 2nd Micro-Action Analysis Grand Challenge in Conjunction with ACM MM'25.

Chuang Wang, Weidong Chen 0010, Xu Cui, Yiming Zhao, Zhaobo Qi, Pengqi Huang, Xinyan Liu 0008, Weigang Zhang

In contrast to traditional action recognition, Micro-Action Recognition focuses on identifying subtle, low-amplitude movements, which was constrained by two kinds of challenges. The first challenge is the spatial imbalance, where small, critical action regions are easily overwhelmed by vast, irrelevant backgrounds, leading to a low signal-to-noise ratio. The second challenge is the class distribution imbalance, where the natural occurrence of actions follows a long-tailed distribution, causing models to be biased towards common actions. To address these specific issues, our framework introduces two targeted solutions. To mitigate spatial imbalance, a YOLOv12-based detection module has been used to localize and crop salient body parts, forcing the model to focus on action-relevant regions. Concurrently, to mitigate class imbalance, this study implement a dynamic oversampling strategy combined with temporal data augmentation, effectively re-weighting the training process to improve performance on rare categories. Integrated with a V-JEPA2 backbone and a multi-classifier ensemble, our approach demonstrates its efficacy by securing second place in the ACM MM'25 Micro-Action Analysis Challenge with an F1-score of 76.98%.

Qiankun Li, Qiupu Chen, Huabao Chen, Feng He, Depeng Li 0001, Zhigang Zeng

Recent advances in video action recognition have achieved remarkable performance in coarse-grained macro-action classification by leveraging large-scale visual backbones and transformer architectures. However, extending these successes to fine-grained micro-action recognition remains a fundamental challenge due to the subtlety, brevity, and low motion intensity of micro-actions. In this paper, we propose a high-capacity framework for micro-action recognition, enhancing both representation learning and decision robustness. We scale to large-scale backbones using the VideoMAEv2 Giant model, enabling the extraction of finer spatial-temporal features. A Temporal-Spatial Connector (TSC) is introduced to dynamically highlight discriminative temporal frames and spatial regions, strengthening the model's focus on subtle motion cues critical for micro-action identification. To stabilize optimization and fully exploit the capacity of large models, we design a four-phase progressive training strategy, encompassing linear probing, full fine-tuning, connector-specific optimization, and classifier head refinement. Furthermore, we propose a novel ensemble decision mechanism that integrates Top-K predictions from diverse models via a Large Language Model (LLM), enhancing prediction consistency and robustness through multimodel consensus. Our method achieves an F1mean of 76.54% on the MA-52 dataset, ranking 3rd in the 2025 Micro-Action Analysis Grand Challenge and advancing the state of the art in fine-grained video understanding.

Kun Li 0008, Dan Guo 0001, Xiaobai Li, Haoyu Chen 0001, Pengyu Liu 0005, Fei Wang 0073, Jingjing Hu, Guoying Zhao 0001, Meng Wang 0001

Micro-Actions (MAs) are a crucial form of non-verbal communication in social interactions, with promising applications in human emotion analysis. Although the topic has attracted considerable research interest, progress has been hindered by the lack of publicly available benchmark datasets. To address this gap, the Micro-Action Analysis Grand Challenge (MAC) is organized annually. This paper presents an overview of the 2nd Micro-Action Analysis Grand Challenge, held in conjunction with ACM Multimedia 2025. We provide a comprehensive summary of the challenge, including its dataset, evaluation protocol, results, and discussion. The top-ranked solutions are highlighted to offer valuable insights for researchers, and potential future directions are outlined to guide ongoing developments in this area. The goal of this grand challenge is to foster innovative research in micro-action analysis and advance research in the human-centric action understanding community.

Robin-Nico Kampa, Fabian Deuser, Konrad Habel, Norbert Oswald

Plant phenotyping involves analyzing observable characteristics of plants to better understand their growth, health, and development. In the context of deep learning, this analysis is often approached through single-view classification or regression models. However, these methods often fail to capture all information required for accurate estimation of target phenotypic traits, which can adversely affect plant health assessment and harvest readiness prediction. To address this, the Growth Modelling (GroMo) Grand Challenge at ACM Multimedia 2025 provides a multi-view dataset featuring multiple plants and two tasks: Plant Age Prediction and Leaf Count Estimation. Each plant is photographed from multiple heights and angles, leading to significant overlap and redundancy in the captured information. To learn view-invariant embeddings, we incorporate 24 views, referred to as the selection vector, in a random selection. Our ViewSparsifier approach won both tasks. For further improvement and as a direction for future research, we also experimented with randomized view selection across all five height levels (120 views total), referred to as selection matrices.

Shreya Bansal, Ruchi Bhatt, Amanpreet Chander, Rupinder Kaur, Malya Singh, Mohan Kankanhalli, Abdulmotaleb El Saddik, Mukesh Saini

Understanding plant growth dynamics is a critical component of modern agricultural research, with applications in yield prediction, phenotyping, and sustainable crop management. Despite recent advances in computer vision and deep learning, progress in plant growth modeling has been constrained by the lack of publicly available, high-resolution, multiview, and temporally rich datasets. To address this gap, we introduce Growth Modelling GroMo25, the first international challenge on plant growth modeling using multiview imagery. In this challenge, we propose a dataset that comprises high-resolution images of four crops: wheat, mustard, radish, and okra, captured at consistent time intervals from multiple camera viewpoints under controlled environmental conditions. The challenge focuses on two key tasks: (1) plant age prediction and (2) leaf count estimation, both requiring models to use spatial and temporal plant features. GroMo25 attracted participation from multiple teams worldwide, encouraging benchmarking and innovation in vision-based plant phenotyping. The GitHub repository is publicly available at https://github.com/mriglab/GroMo-Plant-Growth-Modeling-with-Multiview-Images.

Y. Hop Nguyen, Doan Anh Phan Huu, Trung Thai Tran, Nhat Nam Mai, Van Toi Giap, Thao Thi Phuong Dao, Trung-Nghia Le

We present a unified vision-language framework tailored for ENT endoscopy image analysis that simultaneously tackles three clinically-relevant tasks: image classification, image-to-image retrieval, and text-to-image retrieval. Unlike conventional CNN-based pipelines that struggle to capture cross-modal semantics, our approach leverages the CLIP ViT-B/16 backbone and enhances it through Low-Rank Adaptation, multi-level CLS token aggregation, and spherical feature interpolation. These components collectively enable efficient fine-tuning on limited medical data while improving representation diversity and semantic alignment across modalities. To bridge the gap between visual inputs and textual diagnostic context, we introduce class-specific natural language prompts that guide the image encoder through a joint training objective combining supervised classification with contrastive learning. We validated our framework through participation in the ACM MM'25 ENTRep Grand Challenge, achieving 95% accuracy and F1-score in classification, Recall@1 of 0.93 and 0.92 for image-to-image and text-to-image retrieval respectively, and MRR scores of 0.97 and 0.96. Ablation studies demonstrated the incremental benefits of each architectural component, validating the effectiveness of our design for robust multimodal medical understanding in low-resource clinical settings.