Over the last years, 3D morphable models (3DMMs) have emerged as a state-of-the-art methodology for modeling and generating expressive 3D avatars. However, given their reliance on a strict topology, along with their linear nature, they struggle to represent complex full-head shapes. Following the advent of deep implicit functions, we propose imHead, a novel implicit 3DMM that not only models expressive 3D head avatars but also facilitates localized editing of the facial features. Previous methods directly divided the latent space into local components accompanied by an identity encoding to capture the global shape variations, leading to expensive latent sizes. In contrast, we retain a single compact identity space and introduce an intermediate region-specific latent representation to enable local edits. To train imHead, we curate a large-scale dataset of 4K distinct identities, making a step-towards large scale 3D head modeling. Under a series of experiments we demonstrate the expressive power of the proposed model to represent diverse identities and expressions outperforming previous approaches. Additionally, the proposed approach provides an interpretable solution for 3D face manipulation, allowing the user to make localized edits.
论文检索
输入标题、作者或关键词,从 2,072 篇学术成果中精准定位
Anomaly Detection of Integrated Circuits Package Substrates Using the Large Vision Model SAIC: Dataset Construction, Methodology, and Application
PDF ↗Anomaly detection plays a crucial role in the industrial sector, especially in ensuring the quality of integrated circuits (IC), which are critical for product reliability and performance. With increasing demands for higher quality standards, anomaly detection during the IC manufacturing process has become a significant research focus. However, the progress of IC anomaly detection is hampered by the scarcity of defective samples and the shortage of well-defined annotations. To address this challenge, this paper focuses on the research in the field of IC, especially on ceramic package substrates (CPS). We construct a systematic automated optical inspection (AOI) equipment, and based on this, collected large-scale CPS 2D images to build a novel anomaly detection dataset (CPS2D-AD), which offers copious samples with precise annotations, including category, mask, and bounding box. To the best of our knowledge, CPS2D-AD is the largest dataset in the field of IC. Meanwhile, we conduct an extensive benchmark of CPS2D-AD, intending to supplement existing research by providing a baseline for the detection and localization of anomalies in high-resolution data of ceramic package substrates. In addition, we have developed a novel large vision model, Segment Any Integrated Circuits (SAIC), by embedding-based distillation mechanism based on CPS2D-AD datasets. Our CPS2D-AD is the first open-source anomaly detection dataset about ceramic package substrates, which can be accessed at https://github.com/Bingyang0410/CPS2D-AD.
High-quality open-source datasets have emerged as a pivotal catalyst driving the swift advancement of deep learning, while facing the looming threat of potential exploitation. Protecting these datasets is of paramount importance for the interests of their owners. The verification of dataset ownership has evolved into a crucial approach in this domain; however, existing verification techniques are predominantly tailored to supervised models and contrastive pre-trained models, rendering them ill-suited for direct application to the increasingly prevalent masked models. In this work, we introduce the inaugural methodology addressing this critical, yet unresolved challenge, termed Dataset Ownership Verification for Masked Modeling (DOV4MM). The central objective is to ascertain whether a suspicious black-box model has been pre-trained on a particular unlabeled dataset, thereby assisting dataset owners in safeguarding their rights. DOV4MM is grounded in our empirical observation that when a model is pre-trained on the target dataset, the difficulty of reconstructing masked information within the embedding space exhibits a marked contrast to models not pre-trained on that dataset. We validated the efficacy of DOV4MM through ten masked image models on ImageNet-1K and four masked language models on WikiText-103. The results demonstrate that DOV4MM rejects the null hypothesis, with a p-value considerably below 0.05, surpassing all prior approaches. Code is available at https://github.com/xieyc99/DOV4MM.
Traditional alignment methods for Large Vision and Language Models (LVLMs) primarily rely on human-curated preference data. Human-generated preference data is costly; machine-generated preference data is limited in quality; and self-supervised preference data often introduces hallucinations. To overcome these limitations, we propose a novel Panel-of-Peers learning framework inspired by collaborative learning among humans. This approach leverages a panel of LVLMs, each evaluating and learning from their collective outputs through an iterative self-improvement process. By simulating a peer review system, our models generate, assess, and refine outputs in response to a curated set of prompts, mimicking a classroom learning environment. We demonstrate that this methodology enhances model performance without requiring extensive human-labeled datasets. Our experiments show significant improvement across multiple benchmarks, demonstrating the potential of peer evaluations as a scalable alternative to self-supervised alignment. Notably, we show that Panel-of-Peers increases the average score on fifteen benchmarks from 48% to 57%.
Recent advancements in Multimodal Large Language Models (MLLMs), coupled with the progress of reinforcement learning, have substantially enhanced reasoning and decision-making across modalities, including text, vision, audio, and video. This tutorial introduces the fundamental principles, methodologies, and practical applications of MLLM reasoning, with a particular emphasis on strengthening reasoning capabilities in multilingual and cross-domain settings. We further discuss the key challenges and limitations of current multimodal reasoning approaches, as well as future directions for advancing the field. By highlighting how MLLMs support enhanced reasoning and planning in cross-lingual and cross-domain contexts, this session aims to equip researchers and practitioners with the conceptual foundations and practical tools needed to effectively integrate MLLM reasoning into their work.
Embodied intelligence, evolving from rule-based control to learning-driven systems, has primarily focused on rigid robots, whose limitations in flexibility and adaptability have spurred research into soft-bodied platforms inspired by biological organisms. Soft robots offer adaptive, safe solutions for human collaboration and complex environments but face challenges from underactuation and nonlinear dynamics. This workshop centers on multimodal perception and decision-making in soft robotics, gathering researchers to explore cutting-edge technologies, challenges, and solutions across areas like embodied navigation, manipulation, and control, with presentations and discussions on novel findings and methodologies.
Multimodal human understanding is an evolving interdisciplinary field integrating computer science, psychology, and social sciences to model human perception, behaviour, and biases in multimodal data. While recent advancements in multimodal learning excel in tasks like image-text synthesis, they often overlook nuanced human-centric dynamics---such as cultural, political, and individual influences on how modalities (e.g., text and images) interact, complement, or contradict each other. The 4th International Workshop on Multimodal Human Understanding (MUWS) aims at addressing these challenges, fostering novel solutions that explicitly model human perception, behaviour, and biases in multimodal data, with a particular emphasis on real-world challenges in web and social media analysis. This year edition covers two tracks: (1) human-centred multimodal understanding, such as quantifying social biases, analysing sentiment and hate speech, and modelling cross-modal interactions through interdisciplinary theories (e.g., semiotics, gestalt psychology); and (2) Multimodal understanding of global events, supported by a newly curated dataset covering news articles with diverse stances, which facilitates research on cultural framing, societal impact, and bias mitigation in vision-language models. The event features two keynotes from renowned experts from journalism and computer science, research presentations for six accepted papers, and interactive discussions to explore and discuss cutting-edge methodologies and applications in multimodal human understanding. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3728481
Imbuing embodied agents with non-verbal behavior offers significant benefits for agent-human interactions. Despite extensive research on the creation of non-verbal behaviors, the field lacks a standardized benchmarking practice. Researchers rarely compare their findings with previous studies, and when they do, the comparisons are often not methodologically aligned. The GENEA Workshop 2025 aims to bring together the non-verbal behavior generation community to discuss major challenges and solutions in the field and determine the most effective ways to advance it. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746268.
The Second ACM Workshop on AI-Powered Question & Answering Systems for Multimedia (AIQAM'25) was held on 27 October 2025 in Dublin, Ireland, co-located with ACM Multimedia 2025. The workshop's main objective is to create a collaborative and inclusive space for researchers at the intersection of Artificial Inteligence (AI), large language models (LLMs), multimodal information retrieval, and question answering systems. Building on the success of its first edition (AIQAM'24) at ICMR 2024, AIQAM'25 provided a forum for presenting novel methods, applications, and surveys, which address the challenges of integrating text, image, audio, and video data into QA systems. The programme featured contributions ranging from methodological advances in reasoning and evaluation frameworks, domain-specific applications in education and sustainability, to a survey work reviewing the state of multimedia retrieval-augmented QA. This summary paper outlines the objectives and scope of the workshop, describes its format, highlights the keynote and accepted contributions, and acknowledges the efforts of the organising committee.
SUMAC 2025 is the 7th edition of the workshop on analySis, Understanding and proMotion of heritAge Contents. It is held in Dublin, Ireland, on 27 October and is co-located with the 33rd ACM International Conference on Multimedia. The workshop's objective is to present and discuss the latest and most significant trends, challenges, and advances in the fields of machine learning, signal processing, multimodal techniques, and human-machine interaction. The workshop is dedicated to the valorization of cultural heritage, with an emphasis on unlocking and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.
Reliable detection of conversational errors and user-initiated corrections is critical for effective human-robot interaction (HRI). In this study, we present a comprehensive multimodal approach leveraging temporal window processing, targeted feature engineering, and a MiniRocket + Ridge classification pipeline to address the challenges introduced by the ERR@HRI 2.0 dataset. Our methodology systematically integrates multimodal data streams, including facial expressions, acoustic features, and linguistic embeddings, to predict robot failures and user reactions. Experimental results demonstrate significant improvements over baseline models in event-level detection performance. Notably, linguistic features derived from transcript embeddings emerged as the most informative modality, substantially enhancing model performance. However, we observed challenges associated with managing false positives at the event level, suggesting avenues for future refinement in adaptive thresholding and sequential post-processing techniques. Our findings underscore the importance of careful feature selection and robust temporal modelling in developing effective real-time error detection systems for conversational robots. Our code is available online. https://github.com/Ruddy202/err-hri-2.0-armas.git.
The pervasive spread of multimedia misinformation and disinformation presents a significant challenge to information integrity, demanding robust and efficient verification methodologies. Accurately assessing the authenticity and context of complex multimedia content across diverse platforms requires advanced analytical capabilities. This work introduces Ægis, a novel AI-enhanced solution developed as the authors' submission to the ACMMM'25 - Grand Challenge on Multimedia Verification, aimed at improving the efficiency of multimedia verification. The proposed solution integrates both state-of-the-art and traditional verification tools to comprehensively address various aspects of verification tasks, including event summarization, forensic analysis, and evidence validation. In addition to selectively applied verification modules, large language models (LLMs) are leveraged extensively to fuse findings, perform reasoning, and generate structured verification reports-significantly streamlining the verification process. Demonstrated on the competition multimedia dataset, Ægis accurately validates content integrity, extracts geospatial and temporal information, and identifies content origin across platforms, offering reliable and ethical support for real-world fact-checking.
Asynchronous Video Interviews (AVIs) allow candidates to record responses to predefined questions using digital devices, offering both flexibility and remote accessibility. Assessing personality traits and interview performance via AVIs provides organizations with valuable insights into candidate profiles and facilitates the prediction of future job performance. However, prior benchmark challenges, whose datasets were predominantly sourced from social media, suffer from suboptimal construct and methodological validity, limiting their utility for model development and real-world applications. To address these limitations, we introduce the AVI Grand Challenge at ACM Multimedia 2025, featuring a novel dataset of mock AVIs comprising 3,876 videos from 646 participants in a simulated job application procedure. Interview questions were carefully designed to reflect real-world selection contexts and elicit personality expressions grounded in Trait Activation Theory. Personality traits and job competencies were annotated by trained evaluators and professional recruiters, ensuring both methodological rigor and ecological validity. The solutions and algorithms developed in this challenge are analyzed and summarized in this paper to foster the development of fair, reliable, and AI-driven hiring assessments.
The rapid proliferation of AI-generated media, particularly hyper-realistic deepfakes, has underscored the critical need for robust detection systems to mitigate risks such as misinformation and identity theft. However, state-of-the-art deepfake detectors remain vulnerable to adversarial attacks-subtle perturbations designed to evade classification. To address this gap, we organized the Adversarial Attacks on Deepfake Detectors (AADD-2025) challenge, a competitive evaluation aimed at advancing methodologies to expose and strengthen weaknesses in deepfake detection models. The challenge tasked participants with generating adversarial examples capable of evading four diverse classifiers (including ResNet, DenseNet, and two blind models) while preserving structural similarity to original deepfakes. A dataset comprising 16 subsets of high- and low-quality deepfake images generated by GAN-based and diffusion models (e.g., StableDiffusion, StyleGAN3) was provided. Participants were evaluated using a weighted combination of Structural Similarity Index (SSIM) and attack success rates across all classifiers. Thirteen teams proposed innovative solutions leveraging techniques such as latent-space manipulation, ensemble gradient optimization, surrogate modeling, and frequency-domain perturbation. Top-performing approaches, including MR-CAS (1st place), Safe AI (2nd place), and RoMa (3rd place), achieved high SSIM scores (0.74-0.93) while successfully misleading classifiers. Notably, MR-CAS's latent diffusion model inversion strategy and Safe AI's consensus-orthogonal gradient weighting framework demonstrated superior transferability across architectures, including Vision Transformers. The challenge revealed critical insights: latent-space attacks outperformed pixel-level methods, ensemble-based strategies enhanced cross-model robustness, and adversarial perturbations optimized for both CNNs and transformers proved most effective. However, gaps persist in generalizing attacks across heterogeneous models and maintaining perceptual fidelity, highlighting the urgency of developing adaptive defenses and hybrid detection mechanisms. By fostering collaboration and innovation, AADD-2025 provides a benchmark for evaluating adversarial robustness in deepfake detection and underscores the need for resilient systems in the era of AI-generated media.
Nüshu (Jiangyong Nüshu), a script developed and practiced exclusively by women in China, holds recognition as National Intangible Cultural Heritage. This paper presents the design and implementation of an immersive multimodal interaction system centered on The Song of Nüshu, a foundational cultural artifact. By integrating Natural Interaction (NI) and Mixed Reality (MR) technologies, the system constructs a visual environment inspired by classical Chinese landscape painting aesthetics. Within this experiential space, users engage with representations of pivotal female historical figures and events interwoven with Nüshu textual elements. This framework enables participant-driven narrative construction that foregrounds Nüshu's aesthetic dimensions and socio-cultural significance. Our work advances a heritage preservation methodology that innovatively reinterprets tradition while explicitly foregrounding women's historical agency and expressive practices.
Recent advances in AI-generated content have fueled the rise of highly realistic synthetic videos, posing severe risks to societal trust and digital integrity. Existing benchmarks for video authenticity detection typically suffer from limited realism, insufficient scale, and inadequate complexity, failing to effectively evaluate modern vision-language models against sophisticated forgeries. To address this critical gap, we introduce AEGIS, a novel large-scale benchmark explicitly targeting the detection of hyper-realistic and semantically nuanced AI-generated videos. AEGIS comprises over 10,000 rigorously curated real and synthetic videos generated by diverse, state-of-the-art generative models, including Stable Video Diffusion, CogVideoX-5B, KLing, and Sora, encompassing open-source and proprietary architectures. In particular, AEGIS features specially constructed challenging subsets enhanced with GPT-4o-refined prompts, creating unprecedentedly realistic scenarios for rigorous robustness evaluation. Furthermore, we provide multimodal annotations spanning Semantic-Authenticity Descriptions, Motion Features, and Low-level Visual Features, facilitating authenticity detection and supporting downstream tasks such as multimodal fusion and forgery localization. Extensive experiments using advanced vision-language models demonstrate limited detection capabilities on the most challenging subsets of AEGIS, highlighting the dataset's unique complexity and realism beyond the current generalization capabilities of existing models. In essence, AEGIS establishes an indispensable evaluation benchmark, fundamentally advancing research toward developing genuinely robust, reliable, and broadly generalizable video authenticity detection methodologies capable of addressing real-world forgery threats. Our dataset is avaliable on https://huggingface.co/datasets/Clarifiedfish/AEGIS.
Video content comprehension is essential for various applications, ranging from video analysis to interactive systems. Despite advancements in large-scale vision-language models (VLMs), these models often struggle to capture the nuanced, spatiotemporal details essential for thorough video analysis. To address this gap, we introduce Video-CoT, a groundbreaking dataset designed to enhance spatiotemporal understanding using Chain-of-Thought(CoT) methodologies. Video-CoT contains 192,000 fine-grained spatiotemporal question-answer pairs and 23,000 high-quality CoT-annotated samples, providing a solid foundation for evaluating spatiotemporal understanding in video comprehension. Addition- ally, we provide a comprehensive benchmark for assessing these tasks, with each task featuring 750 images and tailored evaluation metrics. Our extensive experiments reveal that current VLMs face significant challenges in achieving satisfactory performance, high- lighting the difficulties of effective spatiotemporal understanding. Overall, the Video-CoT dataset and benchmark open new avenues for research in multimedia understanding and support future innovations in intelligent systems requiring advanced video analysis capabilities. By making these resources publicly available, we aim to encourage further exploration in this critical area. Project website: https://video-cot.github.io/ .
Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training data with diverse textual metadata, VLMs acquire open-vocabulary capabilities, solving tasks beyond their training scope. This paper investigates the temporal awareness of VLMs, assessing their ability to position visual content in time. We introduce TIME10k, a benchmark dataset of over 10,000 images with temporal ground truth, and evaluate the time-awareness of 37 VLMs by a novel methodology. Our investigation reveals that temporal information is structured along a low-dimensional, non-linear manifold in the VLM embedding space. Based on this insight, we propose methods to derive an explicit ''timeline'' representation from the embedding space. These representations model time and its chronological progression and thereby facilitate temporal reasoning tasks. Our timeline approaches achieve competitive to superior accuracy compared to a prompt-based baseline while being computationally efficient. All code and data are available at https://tekayanidham.github.io/timeline-page/.
Recent text-to-image generative models facilitate creating vivid images with arbitrary contents that are indistinguishable from authentic ones by naked eyes. Despite progress in synthetic image detection, detecting the image from new generators remains challenging. Because advanced generators leave fewer visible forgery traces, while different generative frameworks produce varied forgery patterns. We notice that generative models consistently struggle with fine-detailed content generation, creating abnormal spatial dependencies among neighboring pixels in complex texture regions. In this paper, we propose a methodology of gazing local detail of forgery (GLDF) for generator agnostic synthetic image detection, which identifies prominent spatial dependencies to capture subtle forgery. Concretely, we design frequency-aware correlation discovering (FACD) module to learn dynamic filters by instance-adaptive frequency masking block for identifying prominent spatial deficiencies, which distributed in different spatial positions with various patterns. Furthermore, we introduce the spatial forgery clue distilling module (SFCD) to iteratively aggregate and refine spatial dependencies from different positions by spatial aggregating and prototype global interacting blocks. Extensive experiments demonstrate that GLDF outperforms state-of-the-art methods on detecting synthetic images from different generators.
This paper addresses the critical challenge of detecting codec-based audio deepfakes in multilingual and dynamically evolving adversarial scenarios. While existing detection systems exhibit performance degradation against codec-generated forgeries and unseen linguistic environments, we propose a novel audio deepfake detection framework ''WhiADD'' enhanced by semantic-acoustic fusion and cross-modal generalization. Our methodology introduces three key innovations: (1) The Union CodecFake (UCF) dataset, synthesized by extending the CodecFake generation pipeline to the multilingual Common Voice corpus, significantly expands acoustic diversity with 1.9M samples across varied phonetic, channel, and codec manipulation patterns. (2) A semantic-prompted Whisper architecture that integrates full-transcript linguistic constraints into decoder fine-tuning, enabling detection of semantic inconsistencies. (3) A gated cross-attention mechanism that dynamically fuses multi-source audio features with the proposed model's frozen encoder outputs, enhancing artifact detection through adaptive attention to pre-trained representations. Extensive experiments demonstrate state-of-the-art performance, achieving 0.55% EER on UCF testing data and less than 3% EER in zero-shot cross-lingual detection (German, French, Italian). The framework reduces false negatives by up to 24% compared to conventional models through improved semantic-acoustic alignment. These advancements establish a robust paradigm for combating evolving codec-based forgeries, bridging the critical gap between acoustic feature engineering and semantic coherence analysis in audio forensics.