论文检索

输入标题、作者或关键词,从 12,319 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
12,319篇论文匹配“Datasets and Benchmarks”
第 189 / 616 页

Trong-Thuan Nguyen, Viet-Tham Huynh, Thao Thi Phuong Dao, Mai-Khiem Tran, Ha Nguyen Thi, Tien To Vu Thuy, Uyen Hanh Tran, Tam V. Nguyen 0002, Minh-Triet Tran, Thanh Dinh Le

Automated analysis of endoscopic imagery is a critical yet underdeveloped component of ENT (ear, nose, and throat) care, hindered by variability in devices and operators, subtle and localized findings, and fine-grained distinctions such as laterality and vocal-fold state. In addition to classification, clinicians require reliable retrieval of similar cases, both visually and through concise textual descriptions. These capabilities are rarely supported by existing public benchmarks. To this end, we introduce ENTRep, the ACM Multimedia 2025 Grand Challenge on ENT endoscopy analysis, which integrates fine-grained anatomical classification with image-to-image and text-to-image retrieval under bilingual (Vietnamese and English) clinical supervision. Specifically, the dataset comprises expert-annotated images, labeled for anatomical region and normal or abnormal status, and accompanied by dual-language narrative descriptions. In addition, we define three benchmark tasks, standardize the submission protocol, and evaluate performance on public and private test splits using server-side scoring. Moreover, we report results from the top-performing teams and provide an insightful discussion.

Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira, Wendy Ju, Micol Spitale, Hatice Gunes, Chien-Ming Huang 0001

The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user intent, prematurely interrupting users, or failing to respond altogether. Detecting and addressing these failures is critical for preventing conversational breakdowns, avoiding task disruptions, and sustaining user trust. To tackle this problem, the ERR@HRI 2.0 Challenge provides a multimodal dataset of LLM-powered conversational robot failures during human-robot conversations and encourages researchers to benchmark machine learning models designed to detect robot failures. The dataset includes 16 hours of dyadic human-robot interactions, incorporating facial, speech, and head movement features. Each interaction is annotated with the presence or absence of robot errors from the system perspective, and perceived user intention to correct for a mismatch between robot behavior and user expectation. Participants are invited to form teams and develop machine learning models that detect these failures using multimodal data. Submissions will be evaluated using various performance metrics, including detection accuracy and false positive rate. This challenge represents another key step toward improving failure detection in human-robot interaction through social signal analysis.

Meng Luo 0010, Hao Fei 0001, Bobo Li 0001, Shengqiong Wu, Qian Liu 0012, Soujanya Poria, Erik Cambria, Mong-Li Lee, Wynne Hsu

Understanding fine-grained sentiment dynamics in human conversations is a central goal for next-generation artificial intelligence, especially in scenarios where interactions are rich in both modalities and context. To advance research in this area, we organize the Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) challenge to the community of aspect-based sentiment analysis. The MCABSA challenge introduces two novel subtasks: 1) Panoptic Sentiment Sextuple Extraction, panoramically recognizing holder, target, aspect, opinion, sentiment, and rationale from multi-turn, multi-party multimodal dialogue; and 2) Sentiment Flipping Analysis, detecting the dynamic sentiment transformation throughout the conversation along with the causal reasons. To support these tasks, we present the PanoSent dataset, a high-quality, large-scale benchmark featuring multi-turn, multi-party dialogues annotated with both explicit and implicit sentiment elements across text, image, audio, and video modalities. PanoSent offers extensive real-world scenario coverage, providing a comprehensive testbed for multimodal conversational sentiment analysis. The challenge has attracted widespread participation from both academia and industry, with over 30 teams registered and more than 100 successful submissions. In this paper, we introduce the task, dataset, and evaluation settings, summarize the systems of the top teams, and discuss the findings of the participants. Further details of the challenge can be found at https://panosent.github.io/MM25-challenge.

Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng

With the explosive growth of video-centric social media platforms, understanding and predicting video popularity has become a crucial problem in both academia and industry. This year, the Social Media Prediction (SMP) Challenge expands its scope by introducing a dedicated video track, shifting the focus from static images to dynamic, multimodal video content. We introduce the Social Media Prediction for Videos (SMPV) task and release a large-scale, multimodal benchmark dataset, SMPD-Video, with more than 6K short-form videos, including vision language metadata, user profiles, and popularity labels. This challenge invites global researchers to develop predictive algorithms that integrate spatial-temporal dynamics, multimodal learning, and user-video interactions to forecast video popularity in real-world social temporal streams. With the participation and contribution of top teams around the world, the challenge has seen continuous performance improvements in recent years, driven by technological advancements. SMP Challenge Homepage: www.smp-challenge.com.

Peng Wang 0210, Pujun Xue, Xiaofeng Liu 0006, Tongjuan Ji

Generating diverse and contextually appropriate facial reactions remains a significant challenge due to variability in individual responses, limited explainability, and insufficient modeling of contextual cues. In this study, we propose a multimodal framework that integrates behavioral memory, dynamic attention control, and cognitive style modeling to generate personalized and psychologically grounded facial reactions in dyadic interactions. Our method models the causal link between speaker behavior and listener response by incorporating frame-level behavioral cues, personality traits, and cognitive processing styles. The proposed system consists of three core components: a behavioral memory module that captures temporal context across conversation turns; a Personalized Personality Recognition Style (PPRS) module that infers cognitive tendencies via dual-path learning based on the Big Five personality traits; and a transformer-based generative module equipped with diffusion modeling and context-aware attention gating. This design enables the generation of expressive, individualized responses even during silence or scene transitions. We conduct extensive evaluations on the REACT2025 benchmark using the MARS dataset. Results show that our method outperforms state-of-the-art models in appropriateness (FRCorr ↑0.71), diversity (FRDiv ↑0.1405), and synchrony (FRSyn ↑47.77), ranking 1st in the offline track and 3rd in the online setting. These findings highlight the framework's effectiveness in simulating human-like, emotionally congruent reactions while offering interpretability grounded in personality psychology.

Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cheng Luo, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi 等

In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025

Zizheng Guo 0002, Bochao Zou, Yinuo Jia, Xiangyu Li, Huimin Ma 0001

Micro-expressions (MEs) are involuntary, low-intensity, and short-duration facial expressions that often reveal an individual's genuine thoughts and emotions. Most existing ME analysis methods rely on window-level classification with fixed window sizes and hard decisions, which limits their ability to capture the complex temporal dynamics of MEs. Although recent approaches have adopted video-level regression frameworks to address some of these challenges, interval decoding still depends on manually predefined, window-based methods, leaving the issue only partially mitigated. In this paper, we propose a prior-guided video-level regression method for ME analysis. We introduce a scalable interval selection strategy that comprehensively considers the temporal evolution, duration, and class distribution characteristics of MEs, enabling precise spotting of the onset, apex, and offset phases. In addition, we introduce a synergistic optimization framework, in which the spotting and recognition tasks share parameters except for the classification heads. This fully exploits complementary information, makes more efficient use of limited data, and enhances the model's capability. Extensive experiments on multiple benchmark datasets demonstrate the state-of-the-art performance of our method, with an STRS of 0.0562 on CAS(ME)3 and 0.2000 on SAMMLV. The code is available at https://github.com/zizheng-guo/BoostingVRME.

Tianyi Zhang 0013, Tianhua Qi, Antonis Koutsoumpis, Yuan Zong, Wenming Zheng, Janneke K. Oostrom, Djurre Holtrop, Zhaojie Luo, Reinout E. de Vries

Asynchronous Video Interviews (AVIs) allow candidates to record responses to predefined questions using digital devices, offering both flexibility and remote accessibility. Assessing personality traits and interview performance via AVIs provides organizations with valuable insights into candidate profiles and facilitates the prediction of future job performance. However, prior benchmark challenges, whose datasets were predominantly sourced from social media, suffer from suboptimal construct and methodological validity, limiting their utility for model development and real-world applications. To address these limitations, we introduce the AVI Grand Challenge at ACM Multimedia 2025, featuring a novel dataset of mock AVIs comprising 3,876 videos from 646 participants in a simulated job application procedure. Interview questions were carefully designed to reflect real-world selection contexts and elicit personality expressions grounded in Trait Activation Theory. Personality traits and job competencies were annotated by trained evaluators and professional recruiters, ensuring both methodological rigor and ecological validity. The solutions and algorithms developed in this challenge are analyzed and summarized in this paper to foster the development of fair, reliable, and AI-driven hiring assessments.

Takahiro Komamizu, Marc A. Kastner 0001, Yasutomo Kawanishi, Trung Thanh Nguyen 0006, Junan Chen 0004

The IntentVC Challenge, held in conjunction with ACM Multimedia 2025, introduces a novel benchmark for intention-oriented controllable video captioning. Unlike conventional captioning methods that generate generic, scene-level summaries, IntentVC focuses on intention-specific generation. Participants are required to produce captions explicitly conditioned on user-defined intentions, such as emphasizing a specific object tracked within a video. To support this task, the challenge provides an extended version of the LaSOT dataset annotated with intention-focused captions across 70 object categories. A standardized evaluation protocol and public leaderboard enable fair and reproducible comparison among submitted methods. By advancing research in personalized and adaptive video understanding, IntentVC offers a platform for exploring controllable vision-language modeling with practical relevance for accessibility, retrieval, and human-AI interaction. As a result, a total of 23 teams and 58 active participants have participated, and a total of 1,443 entries have been submitted. More information and resources are available at https://sites.google.com/view/intentvc/.

Dong Chen 0017, Fei Gao, Zhengqing Hu, Xiaojun Chang

We introduce the MIRAGE Challenge, a comprehensive benchmark for multimodal interleaved reasoning and generation, to ACM MM 2025. The challenge aims to evaluate models' abilities to both understand and generate content from complex, multimodal contexts consisting of interlinked images and text. The challenge is accompanied by the MIRAGE Dataset, comprising 263.7K high-quality instruction-response pairs across 35 tasks in two tracks: reasoning and generation. These pairs span 20 diverse scenarios, from surveillance to artistic creation, ensuring broad coverage. The challenge includes seven major categories: Multi-Image Reasoning, Document and Knowledge-Based Understanding, Interactive Multi-Modal Communication, Multi-Image Discrimination, Sequential Visual Generation, Material-based Image Coloring, and Visual Reference Customization. Hosting the MIRAGE Challenge at MM 2025 will drive significant progress in unified multimodal learning and inspire broad involvement in developing more versatile AI systems capable of both understanding and generating multimodal content. Challenge details and participation information are available at https://mm25mirage.github.io/mirage/.

Haoxuan Li, Wei Song, Aofan Liu, Peiwu Qin

Document Visual Question Answering (Document VQA) faces significant challenges when processing long documents in low-resource environments due to context limitations and insufficient training data. This paper presents AdaDocVQA, a unified adaptive framework addressing these challenges through three core innovations: a hybrid text retrieval architecture for effective document segmentation, an intelligent data augmentation pipeline that automatically generates high-quality reasoning question-answer pairs with multi-level verification, and adaptive ensemble inference with dynamic configuration generation and early stopping mechanisms. Experiments on Japanese document VQA benchmarks demonstrate substantial improvements with 83.04% accuracy on Yes/No questions, 52.66% on factual questions, and 44.12% on numerical questions in JDocQA, and 59% accuracy on LAVA dataset. Ablation studies confirm meaningful contributions from each component, and our framework establishes new state-of-the-art results for Japanese document VQA while providing a scalable foundation for other low-resource languages and specialized domains. Our code available at: https://github.com/Haoxuanli-Thu/AdaDocVQA.

Daichi Sato, Duc Minh Vo, Khan Md. Anwarus Salam, Hidenori Shoji, Yuma Matsuoka, Takara Taniguchi, Kaito Baba, Hideki Nakayama

The advent of Large Vision-Language Models (LVLMs) has demonstrated significant capabilities in multimodal understanding. However, their application to complex, multi-page documents, particularly in non-English languages like Japanese, remains a significant challenge due to the scarcity of suitable benchmarks. To address this gap, we organized the ''Large Vision---Language Model Learning and Applications (LAVA) Grand Challenge'' at ACM MultiMedia 2025. We present an overview of the competition. We designed a novel, challenging task: a 10-way multiple-choice Visual Question Answering (VQA) task on multi-page Japanese PDF documents. The task demands that models integrate information across multiple pages, text, and figures. We detail the dataset construction, including an annotation and filtering process designed to ensure questions are visually grounded and non-trivial. We also present the competition results, including an analysis of the leaderboard, and discuss the baseline performance of representative models. The LAVA Grand Challenge highlighted both the current capabilities and limitations of LVLMs in practical document understanding scenarios, thereby stimulating future research and providing a robust benchmark in this important domain.

Yiheng Zhang, Zhaofan Qiu, Qi Cai, Yehao Li, Fuchen Long, Yingwei Pan, Ting Yao 0003, Tao Mei 0001

Recent advancements in multimodal AIGC have enabled impressive text-to-video synthesis, but a critical challenge remains: maintaining consistent identity of key subjects across generated frames. To address this limitation, we introduce the Identity-Preserving Video Generation (IPVG) grand challenge. This challenge aims to propel the field toward more controllable generative models by focusing community efforts on preserving identity during the video generation process. To support these efforts, we publicly release the Identity-Preserving Video Benchmark (VIP-200K), a novel dataset comprising approximately 500,000 video-prompt pairs with 200,000 unique identities, each coupled with a reference identity image. Through this grand challenge and dataset, we provide a fertile ground for developing solutions that lead to more user-steerable video synthesis systems. The challenge homepage is https://hidream-ai.github.io/ipvg-challenge.github.io/.

Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar

The rapid proliferation of AI-generated media, particularly hyper-realistic deepfakes, has underscored the critical need for robust detection systems to mitigate risks such as misinformation and identity theft. However, state-of-the-art deepfake detectors remain vulnerable to adversarial attacks-subtle perturbations designed to evade classification. To address this gap, we organized the Adversarial Attacks on Deepfake Detectors (AADD-2025) challenge, a competitive evaluation aimed at advancing methodologies to expose and strengthen weaknesses in deepfake detection models. The challenge tasked participants with generating adversarial examples capable of evading four diverse classifiers (including ResNet, DenseNet, and two blind models) while preserving structural similarity to original deepfakes. A dataset comprising 16 subsets of high- and low-quality deepfake images generated by GAN-based and diffusion models (e.g., StableDiffusion, StyleGAN3) was provided. Participants were evaluated using a weighted combination of Structural Similarity Index (SSIM) and attack success rates across all classifiers. Thirteen teams proposed innovative solutions leveraging techniques such as latent-space manipulation, ensemble gradient optimization, surrogate modeling, and frequency-domain perturbation. Top-performing approaches, including MR-CAS (1st place), Safe AI (2nd place), and RoMa (3rd place), achieved high SSIM scores (0.74-0.93) while successfully misleading classifiers. Notably, MR-CAS's latent diffusion model inversion strategy and Safe AI's consensus-orthogonal gradient weighting framework demonstrated superior transferability across architectures, including Vision Transformers. The challenge revealed critical insights: latent-space attacks outperformed pixel-level methods, ensemble-based strategies enhanced cross-model robustness, and adversarial perturbations optimized for both CNNs and transformers proved most effective. However, gaps persist in generalizing attacks across heterogeneous models and maintaining perceptual fidelity, highlighting the urgency of developing adaptive defenses and hybrid detection mechanisms. By fostering collaboration and innovation, AADD-2025 provides a benchmark for evaluating adversarial robustness in deepfake detection and underscores the need for resilient systems in the era of AI-generated media.

Zhixi Cai, Kartik Kuckreja, Shreya Ghosh 0001, Akanksha Chuchra, Muhammad Haris Khan, Usman Tariq, Tom Gedeon, Abhinav Dhall

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025.

Huiming Zheng, Linjie Zhou, Wei Gao 0003

With the rapid growth of digital content and the increasing demand for high-resolution displays, the efficient compression of screen content images characterized by text, graphics, and UI elements has become an important research field. This paper introduces a new dataset named SCID-Compress900 specially designed for image compression research. The dataset consists of 900 high-quality screen content images, including 500 4K images and 400 1080P images. All these images are mainly composed of text/graphics, reflecting typical screen content scenarios such as office documents, software interfaces, and presentation slides. The dataset covers a diverse range of content, including various font sizes, graphic styles, and color modes, providing a comprehensive testbed for compression algorithms. To demonstrate the effectiveness of SCID-Compress900, we conduct benchmark tests using several deep learning-based image compression methods commonly employed by researchers. The experimental results show that SCID-Compress900 can well differentiate the performance of different compression algorithms. Compared with existing datasets, SCID-Compress900 offers higher resolution, larger scale, and more targeted content, making it an ideal resource for developing and evaluating advanced image compression algorithms for screen content. This dataset will not only promote the research and development of screen content compression technology but also contribute to the standardization and optimization of compression algorithms in practical applications. The project is available at https://openi.pcl.ac.cn/OpenDatasets/SCID-Compress900.

Zhucun Xue

This paper presents a doctoral research focusing on integrating Retrieval-Augmented Generation (RAG) into video-related multimodal tasks. Existing RAG studies predominantly target text, images, or tabular data, overlooking the unique value of video as a knowledge carrier. We address this gap by: 1) proposing AdaVideoRAG, a framework that adaptively allocates retrieval strategies based on query complexity for long-video understanding; 2) developing REViG (RAG-Enhanced Video Generation) to optimize prompt engineering via retrieved knowledge for controllable video synthesis; 3) constructing the UltraVideo dataset (UHD-4K/8K resolution, 100+ themes, 10 structured captions per video) and HiVU/HiVG benchmarks to evaluate RAG-driven video tasks. Experiments validate the effectiveness of our methods, and we outline future plans to unify video understanding and generation through Agentic RAG for AGI-oriented research.

Alexander Filonenko, Ilya Makarov, Andrey V. Savchenko

FaceCluster is an interactive photo management system that leverages our enhanced KP-RPE face recognition model with Embedding Statistical Regularization to organize personal photo collections automatically. Unlike existing cloud-based systems that raise privacy concerns, FaceCluster operates entirely locally while demonstrating high performance across multiple challenging benchmarks, including IJB-C (97.25% TAR@0.01%), TinyFace (74.14% Rank-1), and AgeDB (97.78% accuracy) when trained on the WebFace4M dataset. The demo showcases real-time face detection, clustering, and organization capabilities through an intuitive web interface, enabling users to effortlessly manage large photo collections with a single-command Docker deployment.

Changsheng Gao, Wei Zhou 0021, Guosheng Lin, Weisi Lin

The widespread deployment of large models in resource-constrained environments has underscored the need for efficient transmission of intermediate feature representations. In this context, feature coding, which compresses features into compact bitstreams, becomes a critical component for scenarios involving feature transmission, storage, and reuse. However, this compression process inevitably introduces semantic degradation that is difficult to quantify with traditional metrics. To address this, we formalize the research problem of Compressed Feature Quality Assessment (CFQA), aiming to evaluate the semantic fidelity of compressed features. To advance CFQA research, we propose the first benchmark dataset, comprising 300 original features and 12000 compressed features derived from three vision tasks and four feature codecs. Task-specific performance degradation is provided as true semantic distortion for evaluating CFQA metrics. We systematically assess three widely used metrics -- MSE, cosine similarity, and Centered Kernel Alignment (CKA) -- in terms of their ability to capture semantic degradation. Our findings demonstrate the representativeness of the proposed dataset while underscoring the need for more sophisticated metrics capable of measuring semantic distortion in compressed features. This work advances the field by establishing a foundational benchmark and providing a critical resource for the community to explore CFQA. To foster further research, we release the dataset and all associated source code at https://github.com/chansongoal/Compressed-Feature-Quality-Assessment.

Bei Yan, Zhiyuan Chen, Yuecong Min, Jie Zhang 0071, Jiahao Wang, Xiaozhen Wang, Shiguang Shan

Despite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, which correspond to faithfulness and factuality hallucinations, respectively. Prior studies primarily evaluate faithfulness hallucination at a rather coarse level (e.g., object-level) and lack fine-grained analysis. Additionally, existing benchmarks often rely on costly manual curation or reused public datasets, raising concerns about scalability and data leakage. To address these limitations, we propose an automated data construction pipeline that produces scalable, controllable, and diverse evaluation data. We also design a hierarchical hallucination induction framework with input perturbations to simulate realistic noisy scenarios. Integrating these designs, we construct SHALE, a Scalable HALlucination Evaluation benchmark designed to assess both faithfulness and factuality hallucinations via a fine-grained hallucination categorization scheme. SHALE comprises over 30K image-instruction pairs spanning 12 representative visual perception aspects for faithfulness and 6 knowledge domains for factuality, considering both clean and noisy scenarios. Extensive experiments on over 20 mainstream LVLMs reveal significant factuality hallucinations and high sensitivity to semantic perturbations.