论文检索

输入标题、作者或关键词,从 2,480 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
2,480篇论文匹配“Speech”
第 43 / 124 页

Sherzod Hakimov, David Semedo, Eric Müller-Budack, Marc A. Kastner 0001, Takahiro Komamizu

Multimodal human understanding is an evolving interdisciplinary field integrating computer science, psychology, and social sciences to model human perception, behaviour, and biases in multimodal data. While recent advancements in multimodal learning excel in tasks like image-text synthesis, they often overlook nuanced human-centric dynamics---such as cultural, political, and individual influences on how modalities (e.g., text and images) interact, complement, or contradict each other. The 4th International Workshop on Multimodal Human Understanding (MUWS) aims at addressing these challenges, fostering novel solutions that explicitly model human perception, behaviour, and biases in multimodal data, with a particular emphasis on real-world challenges in web and social media analysis. This year edition covers two tracks: (1) human-centred multimodal understanding, such as quantifying social biases, analysing sentiment and hate speech, and modelling cross-modal interactions through interdisciplinary theories (e.g., semiotics, gestalt psychology); and (2) Multimodal understanding of global events, supported by a newly curated dataset covering news articles with diverse stances, which facilitates research on cultural framing, societal impact, and bias mitigation in vision-language models. The event features two keynotes from renowned experts from journalism and computer science, research presentations for six accepted papers, and interactive discussions to explore and discuss cutting-edge methodologies and applications in multimodal human understanding. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3728481

Amit Kumar Jaiswal 0001, Thomas Mandl 0001, Gautam Kishore Shahi, Durgesh Nandini, Haiming Liu 0002

With the advancement of digital technologies and gadgets, online content has become easily accessible. At the same time, harmful content also spread widely. There are different harmful content types present on various platforms in multiple languages. The topic of harmful content is broad and covers multiple research directions. Users of platforms are affected by all of them. In research, the different forms are mostly analysed separately, e.g. misinformation, cyber-bullying and hate speech. Most research has been conducted for only one platform, for a monolingual situation or on a particular issue. Counter-measures like blocking are down-ranking can make harmful content spreaders to switch platforms and languages to continuously reach a user base. Harmful content does not only appear on social media but also on news media. Spreader share harmful content in posts, news articles, comments and hyperlinks. There is a great need to study harmful content across platforms, languages, and topics. We plan to bring the research on harmful content under one umbrella such that different approaches and novel methods can be shared. The workshop will also cover the currently ongoing issues of war and elections. We propose the workshop, DHOW: Diffusion of Harmful Content on Online Web, which brings together the research on different topics of harmful content. We expect to discuss innovative research work and future research directions. The proposed workshop is the next iteration of DHOW 2024. https://dhow-workshop.github.io previously organized at ACM WebSci 2024 in Stuttgart, Germany.

Omar Shahbaz Khan, Ujjwal Sharma 0001, Gonçalo Marcelino, Aaron Duane, Stevan Rudinac, Marcel Worring, Björn Þór Jónsson 0001

We introduce multiXview, an interactive retrieval framework for synchronized multi-camera video collections. It features a multi-index search engine that supports natural-language queries over visual embeddings, speech transcripts, and scene descriptions. It supports a synchronized multi-stream player offering parallel playback, and a timeline-based navigation view for temporal scoping and faceted exploration. These components address the redundancy and fragmentation of overlapping egocentric and exocentric video feeds and enable users to locate, aggregate, and reconstruct events across partial perspectives. This paper focuses on system design and implementation, with quantitative and qualitative evaluation to take place at the CASTLE 2025 Grand Challenge Interactive Track.

Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira, Wendy Ju, Micol Spitale, Hatice Gunes, Chien-Ming Huang 0001

The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user intent, prematurely interrupting users, or failing to respond altogether. Detecting and addressing these failures is critical for preventing conversational breakdowns, avoiding task disruptions, and sustaining user trust. To tackle this problem, the ERR@HRI 2.0 Challenge provides a multimodal dataset of LLM-powered conversational robot failures during human-robot conversations and encourages researchers to benchmark machine learning models designed to detect robot failures. The dataset includes 16 hours of dyadic human-robot interactions, incorporating facial, speech, and head movement features. Each interaction is annotated with the presence or absence of robot errors from the system perspective, and perceived user intention to correct for a mismatch between robot behavior and user expectation. Participants are invited to form teams and develop machine learning models that detect these failures using multimodal data. Submissions will be evaluated using various performance metrics, including detection accuracy and false positive rate. This challenge represents another key step toward improving failure detection in human-robot interaction through social signal analysis.

Keqi Chen, Wenxin Fu, Qihang Lu, Zekai Sun, Yizhong Geng, Yi Liu, Puyuan Guo, Yingming Gao, Ya Li 0001

Emotional Support Chatbots could unlock potential by providing scalable, low-cost, and personal emotional support, overcoming critical accessibility barriers inherent in traditional counseling. However, current text-based Chatbots fall short in conveying the multimodal empathy crucial in counseling. Humans naturally prefer face-to-face communication with peers to share feelings, encompassing spoken tone, micro-expressions, and body language to convey empathy. To bridge this gap, we propose EMO-Avatar, an LLM-agent-orchestrated framework that integrates emotional reasoning capabilities and multimodal expression in counseling. Our approach introduces two innovations: (1) a Multimodal Emotional Support Agent. EMO-Avatar can follow adaptive instruction across TTS, pose, micro-expressions, and body actions, leading to the generation of highly expressive human animations. (2) a Comforting-Exploration-Action support strategy; EMO-Avatar systematically integrates Hill's three-stage counseling theory into its emotional reasoning capability. Guided by the LLM's reasoning, this strategy informs response generation and displays stage-specific preferences for speech, body language, and expressions. EMO-Avatar can provide deeper emotional support and therapeutic human-like interactions. Experimental validation on the AvaMERG Challenge demonstrates EMO-Avatar's superior performance, achieving top-2 ranking among 20 participants across response appropriateness, multimodal consistency, naturalness, and emotional expressiveness metrics. Our demo is available at https://ai4ai.anonymous-demo.fun/.

Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Li Huang, Haifeng Hu 0001, Yap-Peng Tan

Multimodal Empathetic Response Generation (MERG) is crucial for building emotionally intelligent human-computer interactions. Although large language models (LLMs) have improved text-based ERG, challenges remain in handling multimodal emotional content and maintaining identity consistency. Thus, we propose E3RG, an Explicit Emotion-driven Empathetic Response Generation System based on multimodal LLMs which decomposes MERG task into three parts: multimodal empathy understanding, empathy memory retrieval, and multimodal response generation. By integrating advanced expressive speech and video generative models, E3RG delivers natural, emotionally rich, and identity-consistent responses without extra training. Experiments validate the superiority of our system on both zero-shot and few-shot settings, securing Top-1 position in the Avatar-based Multimodal Empathy Challenge on ACM MM'25. Our code is available at https://github.com/RH-Lin/E3RG.

Han Zhang 0035, Hao Fei 0001, Hong Han 0001, Lizi Liao, Erik Cambria, Min Zhang 0005

The Ava-MERG Challenge at ACM Multimedia 2025 explores avatar-based multimodal empathetic response generation by advancing dialogue systems that respond with emotional awareness across text, speech, and talking-face video. To support this goal, we constructed a high-quality dataset featuring aligned multimodal responses in diverse conversational scenarios. The challenge comprises two progressively complex subtasks centering around multimodal-aware empathetic response generation. We aim to promote research in empathetic AI and affective computing.

Ivan Kukanov, Jun Wah Ng

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and localizing deepfakes, even under novel, unseen attack scenarios. Current state-of-the-art deepfake detectors, while accurate, are often computationally expensive and struggle to generalize to novel manipulation techniques. To address these challenges, we propose multimodal approaches for the AV-Deepfake1M 2025 challenge. For the visual modality, we leverage handcrafted features to improve interpretability and adaptability. For the audio modality, we adapt a self-supervised learning (SSL) backbone coupled with graph attention networks to capture rich audio representations, improving detection robustness. Our approach strikes a balance between performance and real-world deployment, focusing on resilience and potential interpretability. On the AV-Deepfake1M++ dataset, our multimodal system achieves AUC of 92.78% for deepfake classification task and IoU of 0.3536 for temporal localization using only the audio modality.

Zhixi Cai, Kartik Kuckreja, Shreya Ghosh 0001, Akanksha Chuchra, Muhammad Haris Khan, Usman Tariq, Tom Gedeon, Abhinav Dhall

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025.

Jinzhao Zhou, Daniel Leong, Zehong Cao, Thomas Do, Sheng-Fu Liang, Tzyy-Ping Jung, Chin-Teng Lin

This paper presents MindSpeak, a real-time brain-computer interface (BCI) system for recording, processing, and decoding silent speech to enable online multimodal communication between the human brain and a computer, involving both noninvasive multichannel EEG signals and text output. To enable hand-free and brain-only control, our system incorporates steady-state visual evoked potential (SSVEP) for users to select incomplete sentences from a predefined pool and confirm the correctness of decoded words. An intuitive graphical interface is designed for natural communication. We evaluate the effectiveness of our real-time BCI system, which achieves 77.3% accuracy in decoding silent speech and 98.9% accuracy in SSVEP-based selection and confirmation of correct sentences. Unlike existing BCI systems, the presented MindSpeak system significantly expands the application scope of existing BCI systems by enabling users to express complete thoughts through a fully BCI-controlled interactive interface. Our demonstration video is on: https://youtu.be/B1wt1dmCCrg.

Nils Riekers, Marten Risius, Tong Chen 0005

Multimodal hate speech detection targets offensive content expressed through combinations of modalities such as text and images, which often evade detection when analyzed separately. We introduce MAXplain, an interactive framework that addresses both issues via a configurable LLM-based multi-agent architecture. Specialized agents handle distinct subtasks and exchange information through structured dialogues, enabling intrinsic explainability and improved accuracy. The web interface supports human-in-the-loop interaction, including real-time adjustment of agent behaviors and evaluation rules. A browser plugin enables direct inspection of online content. While demonstrated for hate speech detection, MAXplain also supports rapid prototyping for other multimodal tasks.

Andreas Babic, Xihui Chen, Djordje Slijepcevic, Adrian Jaques Böck, Matthias Zeppelzauer

We present CounterHelp, a mobile app designed to support adolescents and young people in generating customized counter speech in response to hateful comments on TikTok. CounterHelp allows users to specify counter strategies to combat hate speech. After users share a hateful comment with CounterHelp, it retrieves TikTok metadata to capture its context. Leveraging large language models, CounterHelp generates customised and context-sensitive counter speech in one of four predefined styles, particularly tailored for adolescents. A user experience lab study confirms the effectiveness and usability of CounterHelp and the generated counter speech.

Han Wang 0053, Zhuoran Wang, Roy Ka-Wei Lee

Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level and segment-level annotations, comprising over 11,714 segments labeled as Normal or across five Offensive categories: Hateful, Insulting, Sexual, Violence, Self-Harm, along with explicit target victim labels. Our three-stage annotation process yields high inter-annotator agreement (Krippendorff's alpha = 0.817). We propose three tasks to benchmark performance: (1) Trimmed Hateful Video Classification, (2) Temporal Hateful Video Localization, and (3) Online Hateful Video Classification. Results highlight substantial gaps in current models, emphasizing the need for more sophisticated multimodal and temporally aware approaches. The HateClipSeg dataset are publicly available at https://github.com/Social-AI-Studio/HateClipSeg.git.

Chenxi Wang, Yusheng Dai, Lei Sun 0010, Jun Du 0002, Jianqing Gao

Recent rapid progress in Text-to-Audio (T2A) models contrasts sharply with the stagnation observed in the evolution of corresponding evaluation benchmarks. Existing benchmarks, such as AudioCaps, suffer from limited diversity and quality, as well as biased category distributions, leading to increasingly questionable reliability in assessing advanced T2A models. This paper introduces AudioAtlas, a comprehensive and balanced evaluation benchmark specifically designed for evaluating T2A models aimed at movie production. Based on an object-centric audio category system, AudioAtlas provides high-quality reference samples characterized by categorical balance and diversity. It includes detailed overall and event-level captions with rich descriptors, plus fine-grained temporal annotations from human experts, enabling thorough evaluation of temporal alignment and semantic accuracy. To enable precise evaluation of temporally-aligned generation across universal categories, two novel metrics are proposed leveraging recent advancements in large-scale Audio Language Models (AudioLLMs) and contrastive learning models. By re-benchmarking six currently influential T2A models, AudioAtlas provides evaluations better aligned with aesthetic considerations, offering clearer optimization directions for movie-production-oriented T2A systems. Additionally, we conduct a comprehensive comparative analysis on temporally-controllable T2A methods with training-based, and promising training-free approaches inspired by region-controllable image generation, clarifying current limitations and pointing out directions for future research. Audio specifically refers to sound event excluding speech and music. Further details are available on the project page: https://audioatlas.github.io/AudioAtlas/

Yuchen Zhang, Tailin Chen, Jiangbei Yue, Yueming Sun, Rahul Singh, Jianbo Jiao, Zeyu Fu

Hate speech poses a persistent threat to society, causing profound harm to both individuals and communities. Detecting such content is essential for promoting safer and more inclusive environments. While previous research has primarily focused on text-based or image-based hate speech detection, video-based hate detection remains relatively underexplored. A key barrier is the limited availability of high-quality video datasets. Existing hateful video datasets are typically limited in scale, diversity, and annotation depth, often labeling hateful content without further distinguishing between explicit and implicit forms. In this work, we present DeHate, which, to the best of our knowledge, is the largest hateful video dataset to date. DeHate comprises 6689 videos collected from two platforms and spanning six social groups. Each video is annotated with fine-grained labels that differentiate explicit, implicit, and non-hateful content, along with segment-level localization of hate, identification of contributing modalities, and specification of the targeted groups. Through detailed analysis of annotated videos across platforms, we reveal distinct patterns in how hateful content is conveyed, offering a comprehensive comparison between explicit and implicit hate in terms of their prevalence and characteristics. Furthermore, we benchmark state-of-the-art models, including both uni-modal and multi-modal architectures, and identify persistent challenges in detecting subtle and context-dependent forms of hate. Our findings highlight the importance of holistic and fine-grained hateful video datasets for advancing research in hate speech detection. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.

Alessandro Ragano, Carl Timothy Tolentino, Kata Szita, Dan Barry, Davoud Shariat Panah, Niall Murray, Andrew Hines

Although audio-augmented reality (AAR) has known applications in music, the use of wearables such as augmented reality (AR) glasses for egocentric audio data capture for music has not been investigated. Current egocentric datasets are mostly focused on speech research, neglecting music's unique demands for tasks such as real-time optimisation or assistive listening. This paper introduces EgoMusic, a multimodal dataset featuring synchronised egocentric audio-visual data captured with AR glasses during live performances, alongside studio-quality audio references. We investigate AR glasses' utility for music and baseline artificial intelligence (AI) approaches for hearing enhancement, positioning EgoMusic as the first dataset that enables research for egocentric music AAR.

Zihao Wang, Shulei Ji, Le Ma 0002, Yuhang Jin, Shun Lei, Jianyi Chen, Haoying Fu, Roger B. Dannenberg, Kejun Zhang

Singing accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal-instrumental separation process. Additionally, they often lack regional accent annotations. To address this, we introduce the Multi-Accent Mandarin Dry-Vocal Singing Dataset (MADVSD). MADVSD comprises over 670 hours of dry vocal recordings from 4,026 native Mandarin speakers across nine distinct Chinese regions. In addition to each participant recording audio of three popular songs in their native accent, they also recorded phonetic exercises covering all Mandarin vowels and a full octave range. We validated MADVSD through benchmark experiments in singing accent recognition, demonstrating its utility for evaluating state-of-the-art speech models in singing contexts. Furthermore, we explored dialectal influences on singing accent and analyzed the role of vowels in accentual variations, leveraging MADVSD's unique phonetic exercises.

Yulong Li 0002, Yuxuan Zhang, Rui Chen, Feilong Tang, Zhixiang Lu, Ming Hu, Jianghao Wu 0001, Haochen Xue, Mian Zhou, Chong Li 等

Artificial intelligence will not achieve genuine empathy until models can reason about the causes of human emotions rather than only label them. Current datasets fail to support this objective, as existing emotional causality datasets primarily focus on textual modalities, lack non-verbal information such as speech and facial expressions, feature relatively short dialogue lengths, and limit research on long-term emotional evolution. Existing annotations concentrate on stimulus-response patterns and lack cross-temporal emotional causal chain annotations, failing to reveal how early events accumulate and ultimately trigger emotional changes. In this work, we introduce Genesis, the first multimodal dialogue dataset supporting long-term emotional causality analysis, which Genesis contains 1,000 dialogues averaging 208 turns each, spanning debate, family, educational, and social scenarios. Through two-layer annotation system: proximal cause identification and long-term causal chain tracking, Genesis labels complex emotional phenomena including cross-modal inconsistencies and long-distance causal dependencies. Our evaluation of 20 mainstream multimodal models reveals limitations in current approaches for long-term emotional causality. We propose Empathica as an evaluation baseline, employing a Recognition-Memory-Attribution architecture that integrates dynamic sliding windows and event aggregation mechanisms to address multimodal emotional causality modeling challenges. Empathica outperforms text-based models GPT-o1, and multimodal model Gemini 1.5 Pro and GPT-4o across all evaluation metrics.

Bingjian Yang, Danni Xu, Kaipeng Niu, Wenxuan Liu 0008, Zheng Wang 0007, Mohan Kankanhalli

The proliferation of online misinformation videos poses serious societal risks. Current datasets and detection methods primarily target binary classification or single-modality localization based on post-processed data, lacking the interpretability needed to counter persuasive misinformation. In this paper, we introduce the task of Grounding Multimodal Misinformation (GroundMM), which verifies multimodal content and localizes misleading segments across modalities. We present the first real-world dataset for this task, GroundLie360, featuring a taxonomy of misinformation types, fine-grained annotations across text, speech, and visuals, and validation with Snopes evidence and annotator reasoning. We also propose a VLM-based, QA-driven baseline, FakeMark, using single and cross-modal cues for effective detection and grounding. Our experiments highlight the challenges of this task and lay a foundation for explainable multimodal misinformation detection. Dataset will be released at https://github.com/yangbingjian/GroundLie360.

Zhuojun Wu, Dong Liu, Juan Liu 0007, Yechen Wang, Linxi Li, Liwei Jin, Hui Bu, Pengyuan Zhang, Ming Li 0026

In natural language communication, emotions are often conveyed through non-verbal sounds (NVs), such as laughter, crying, cough and so on. However, most existing text-to-speech (TTS) corpora lack annotations for these non-verbal sounds, leading to a scarcity of systems capable of generating them. To address this gap, we introduce SMIIP-NV, a non-verbal speech synthesis corpus annotated with both emotions and non-verbal sounds, including laughter, crying, and cough. To the best of our knowledge, SMIIP-NV is the largest publicly available open-source expressive speech corpus that includes non-verbal speech and rich annotations. It comprises 33 hours of speech data, covering five distinct emotions and three types of non-verbal sounds, with detailed transcriptions and precise timestamps for each occurrence of non-verbal sounds. Additionally, the corpus provides annotations for speech segments that contain laughter or crying. To demonstrate the utility of this dataset, we establish a baseline for non-verbal speech synthesis by employing a lightweight large language model (LLM). The SMIIP-NV dataset and static audio demonstrations are publicly available at https://axunyii.github.io/SMIIP-NV. The interactive real-time demonstrations can be accessed at https://huggingface.co/spaces/xunyi/SMIIP-NV_Finetuned_CosyVoice2.