论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 2 / 81 页

Jianquan Liu, Balu Adsumilli, Yukiko Yanagawa, Haiwei Dong 0001

The ACM Multimedia 2025 Industry Program presents a comprehensive overview of how multimodal AI is revolutionizing real-world applications. The program features contributions from over twenty industry leaders, covering a spectrum of domains from content creation and distribution to healthcare, manufacturing, and foundational technology. Keynotes by leaders from NEC and Google DeepMind address critical challenges in business transformation and media integrity, respectively. A dedicated seminar from Google explores advancements in video codecs like AV2 and the development of next-generation quality metrics using Large Language Models (LLMs). The program also includes twelve expert talks and seven demonstrations that showcase practical innovations, such as LLM-driven recommendation systems at Meta, generative AI for industrial optimization by Mitsubishi Electric, and a Siemens-developed protocol for semantic interoperability in manufacturing. These components collectively highlight the profound societal and industrial impact of multimedia research and bridge the gap between academic theory and real-world deployment.

Wei Gao 0003, Ge Li 0002

Different 2D and 3D visual data have been widely used in applications such as UHDTV, mobile phones, autonomous driving, robots, and UAVs, etc. The large-scale data amount has raised the research and development trends of image and video coding, and point cloud coding. Compression technologies can efficiently reliever the burden of communication and storage. This tutorial will discuss diverse aspects of multimedia data compression, including datasets, perception models, deep learning-based coding methods, AI-based standards, open source projects, and future research.

Yicong Li 0004, Junbin Xiao, Angela Yao, Tat-Seng Chua

Video Question Answering (Video QA) has emerged as a central task in multimodal learning. This tutorial provides a comprehensive overview of VideoQA research and highlights new frontiers. We begin with an introduction to VideoQA preliminaries, tracing how methods have adapted from third-person view short videos to capture egocentric and long-ranged spatial-temporal dynamics. We then focus on the impact of large multimodal models. Next, we expand the scope to spatial understanding beyond videos. Each topic is discussed through the lens of tasks, datasets, methods, and evaluation protocols. Finally, we conclude with future directions, including fine-grained and long-ranged video understanding, robustness and trustworthiness, Egocentric and embodied assistance, and omnimodal integration. This tutorial aims to equip participants with both a historical perspective and a forward-looking roadmap for advancing Video QA in the LLM era.

Qiang Sheng 0001, Peng Qi 0005, Tianyun Yang, Yuyan Bu, Wynne Hsu, Mong-Li Lee, Juan Cao 0001

Recent progress of generative AI and the popularity of short-form video-sharing platforms have raised new risks of misinformation video issues, posing a potential threat to online multimedia ecosystems. With the aid of generative AI tools, producing and spreading vivid, persuasive misinformation videos has been easier, while detecting and preventing them has become harder. This tutorial introduces how to characterize, detect, and prevent misinformation videos, which consists of three technical parts: 1) Characterization of AI-generated and human-edited misinformation videos; 2) Detection approaches, covering those tailored for fully generated, manipulated, and human-edited videos; and 3) Prevention strategies, including those effective for the creation and spread phases. This tutorial concludes by discussing the status quo and ongoing challenges and highlighting the promising directions for future research. We expect to bring broader attention to misinformation video issues, gather and communicate with researchers of interest, and facilitate the engagement of those who are new to this field.

Sarmistha Das 0001, Akash Ghosh, Sriparna Saha 0001, Koustava Goswami, K. J. Joseph

Recent advancements in Multimodal Large Language Models (MLLMs), coupled with the progress of reinforcement learning, have substantially enhanced reasoning and decision-making across modalities, including text, vision, audio, and video. This tutorial introduces the fundamental principles, methodologies, and practical applications of MLLM reasoning, with a particular emphasis on strengthening reasoning capabilities in multilingual and cross-domain settings. We further discuss the key challenges and limitations of current multimodal reasoning approaches, as well as future directions for advancing the field. By highlighting how MLLMs support enhanced reasoning and planning in cross-lingual and cross-domain contexts, this session aims to equip researchers and practitioners with the conceptual foundations and practical tools needed to effectively integrate MLLM reasoning into their work.

Wei Zhou 0021, Hadi Amirpour

The rapid expansion of multimedia services, such as video streaming, video conferencing, virtual reality, and cloud gaming, makes maintaining and evaluating high perceptual visual quality essential for user experience and system competitiveness. However, visual content can degrade at multiple stages, including acquisition, compression, transmission, enhancement, and display, where suboptimal enhancement may also introduce artifacts and reduce perceived quality. The core challenge is to reliably measure and predict this perceived quality so that it can be maintained or improved. Perceptual Visual Quality Assessment (PVQA) addresses this by evaluating visual quality from the perspective of human subjects, through subjective studies and objective prediction models. Beyond humans, recent work also extends PVQA to machines and robots, where the goal is to preserve downstream task performance (e.g., segmentation accuracy and planning success) under distortions or bandwidth constraints. This tutorial provides a concise, practice-oriented overview of PVQA: fundamentals and human vision considerations; image and video quality assessment; methods for immersive/3D media; opportunities and challenges in the era of foundation models and GenAI; perceptual optimization loops that close the gap between assessment and decisions in coding, streaming, and embodied perception; and domain applications. Finally, we summarize the key concepts, toolchains, and future opportunities for PVQA to be used in modern multimedia communication.

Siru Zhong, Xixuan Hao, Hao Miao 0001, Yan Zhao 0008, Qingsong Wen, Roger Zimmermann, Yuxuan Liang 0002

Spatio-temporal data mining (STDM) has become crucial in multimedia, driven by the surge of multimodal data from remote sensing, IoT sensors, social media, surveillance systems, mobile devices, and crowdsourced platforms. Traditional single-modal methods, though successful, struggle to capture real-world complexity. Integrating multiple modalities yields richer, more accurate insights, boosting spatio-temporal analysis. This half-day tutorial, MM4ST: Multimodal Learning for STDM, offers a comprehensive overview, covering STDM fundamentals, challenges in aligning and fusing heterogeneous data, advanced multimodal modeling techniques, and emerging research directions. Attendees will acquire practical knowledge to develop scalable and robust spatio-temporal mining solutions. All materials will be publicly available online.

Lipika Dey, Marianna Obrist, Stavroula G. Mougiakakou

MMFood'25, the 1st International Workshop on Multi-modal Food Computing, brings together researchers and practitioners at the intersection of artificial intelligence, computer vision, natural language processing, and sensory modeling to advance the study of food. The workshop highlights how multimodal methods can be applied to food recognition, recommendation, analysis, and monitoring to address pressing challenges in health, nutrition, sustainability, and food culture. It features a rich program including a keynote, an invited talk, paper presentations, a poster session, and a panel discussion on the role of multimodal AI in preserving cultural heritage, fostering sustainable food futures, and enabling personal well-being. By convening experts from academia, industry, and healthcare, MMFood'25 provides a unique platform for fostering interdisciplinary collaboration and for shaping the emerging field of multimodal food computing. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746264.

Irene Viola 0001, Silvia Rossi 0001, Marta Orduna, Maria Torres Vega

Despite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, telesurgery, interconnection in new buildings project chain, cultural heritage, and museum contents communication, are well on their way to integrating immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human-centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks, and end-user infrastructures to enable the next generation of interactive XR applications and services. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746269

Tiesong Zhao, Qian Liu 0001, Zhisheng Yan

Today, truly immersive multimedia systems demand the integration of emerging multi-sensorial media, which go beyond traditional audiovisual signals to include haptics, olfaction, motion capture, electroencephalograms, and other novel media forms. To effectively incorporate these modalities into cutting-edge multimedia systems, advances are needed across the entire pipeline, from processing and encoding to seamless integration. In addition, human-centric factors such as ergonomics and user experience must be considered to ensure practical implementation. Our workshop, the International Workshop on Multi-Sensorial Media and Applications (MSMA'2025), seeks to attract contributions related to multi-sensorial media systems, including system design, evaluation, coding, delivery, media analysis, multi-modal interaction, human factors, ergonomics, and related areas. By fostering collaboration among researchers, MSMA aims to bridge existing work in the field, spark innovation, and push the boundaries of multimedia technology.

Cheng Jin 0001, Mingli Song, Rui Wang 0032, Xingjiao Wu

This workshop addresses next-generation methods in multimedia research, with a focus on content generation, quality assessment, and dataset development. These three areas are foundational for advancing multimedia technologies and applications. Emerging approaches in multimedia content generation, powered by generative AI and multimodal learning, are reshaping domains such as entertainment, advertising, education, and healthcare. At the same time, robust quality assessment is essential to ensure that generated content achieves high standards of perceptual fidelity, semantic consistency, and user satisfaction, thereby determining the real-world impact of multimedia systems. Datasets remain indispensable for training and evaluating algorithms, and innovative strategies in dataset construction-ranging from augmentation and annotation to addressing issues of bias and small-sample imbalance-are driving the development of more reliable and ethical multimedia applications. By convening leading researchers and practitioners, this workshop provides a platform to explore state-of-the-art methods, share best practices, and discuss open challenges in next-generation multimedia research. The goal is to foster interdisciplinary collaboration and inspire innovative solutions that advance the creation, evaluation, and application of multimedia content, setting new benchmarks for the field and shaping the future of multimedia technologies.

Wei Gao 0003, Sam Kwong, Zhu Li 0001, Shan Liu 0001, Ge Li 0002

Point cloud processing and 3D vision have emerged as very hot topics in the multimedia community. Point clouds can give an immersive visual experience and provide accurate structural information of 3D objects and scenes in the applications including virtual reality/augmented reality (VR/AR), autonomous driving, robot navigation, and geo-information systems (GIS). Moreover, 3D Gaussian splatting has become a very powerful and popular tool for 3D reconstruction, generation and rendering, as well as compression, representation and understanding. 3D Gaussian splatting technology can also be deemed as an extension of point cloud processing technology. 3D vision technologies empower the developments of immersive media, embodied artificial intelligence (AI) and unmanned systems. Their challenges in processing, analysis, and applications have attracted significant interest from industry, academia, and standardization bodies. This workshop invites innovative contributions in point cloud processing and 3D Gaussian splatting to propel the advancements of 3D vision technologies.

Hao Fei 0001, Bobo Li 0001, Meng Luo 0010, Qian Liu 0012, Lizi Liao, Fei Li 0021, Min Zhang 0005, Björn W. Schuller, Mong-Li Lee, Erik Cambria

The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing (CogMAEC) was held at ACM Multimedia 2025. It focused on moving emotional AI beyond basic recognition toward deeper cognitive understanding. While traditional multimodal affective computing has emphasized simple emotion detection, the rise of multimodal large language models (MLLMs) has spurred interest in modeling how emotions emerge and evolve in context. The workshop gathered researchers on emotion reasoning, multimodal understanding, and human-computer empathy, exploring how machines can not only recognize emotions but also explain their causes and simulate human-like affective reasoning. The program featured invited talks, oral presentations, and posters spanning perception, interaction, causal modeling, and cognitive grounding. CogMAEC provided a platform to connect researchers across disciplines and foster future work on cognitively aware affective computing. Materials are available at https://CogMAEC.github.io/MM2025.

Zheng Wang 0046, Qianqian Chen, Yiyang Luo, Zhiqiu Ye, Shi Wei, Hanwang Zhang, Tat-Seng Chua

This workshop aims to explore the potential of large generative models to revolutionize the way we interact with multimodal information. A Large Language Model (LLM) represents a sophisticated form of artificial intelligence engineered to comprehend and produce natural language text, exemplified by technologies such as GPT, LLaMA, Flan-T5, ChatGLM, and Qwen, etc. These models undergo training on extensive text datasets, exhibiting commendable attributes including robust language generation, zeroshot transfer capabilities, and In-Context Learning (ICL). With the surge in multimodal content-encompassing images, videos, audio, and 3D models-over the recent period, Large MultiModal Models (LMMs) have seen significant enhancements. These improvements enable the augmentation of conventional LLMs to accommodate multimodal inputs or outputs, as seen in BLIP, Flamingo, KOSMOS, LLaVA, Gemini, GPT-4, etc. Concurrently, certain research initiatives have delved into generating specific modalities, with Kosmos2 and MiniGPT-5 focusing on image generation, and SpeechGPT on speech production. There are also endeavors to integrate LLMs with external tools to achieve a near 'any-to-any' multimodal comprehension and generation capacity, illustrated by projects like Visual-ChatGPT, ViperGPT, MMREACT, HuggingGPT, and AudioGPT. Collectively, these models, spanning not only text and image generation but also other modalities, are referred to as large generative models. This workshop will provide an opportunity for researchers, practitioners, and industry professionals to explore the latest trends and best practices in the field of multimodal applications of large generative models. We also remark that the submissions are not limited to the use of such models. The workshop will also focus on exploring the challenges and opportunities of integrating large language models with other AI technologies such as computer vision and speech recognition. Additionally, the workshop will provide a platform for participants to present their research, share their experiences, and discuss potential collaborations. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3728422

Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar

The proliferation of generative models, particularly Generative Adversarial Networks (GANs) and Diffusion Models, has reshaped multimedia content creation. Alongside creative and commercial opportunities, they have introduced unprecedented risks through the production of highly realistic synthetic content, or deepfakes. These artifacts challenge visual and auditory trust, with major implications for media, security, politics, and law. This workshop provides a forum to examine deepfake technology from forensic, technical, legal, and social perspectives. It will bring together experts to advance robust and explainable detection methods, define benchmarking practices, and address ethical and regulatory frameworks. Topics include detection and attribution, adversarial countermeasures, multimodal analysis, model traceability, legal admissibility of synthetic content, as well as real-world deployment challenges and dataset creation. Further information about the workshop is available at https://iplab.dmi.unict.it/mfs/acm-dff-ws-2025/

Zheng Lian 0004, Shreya Ghosh 0001, Erik Cambria, Zhixi Cai, Guoying Zhao 0001, Abhinav Dhall, Björn W. Schuller, Roland Goecke, Jianhua Tao 0001, Tom Gedeon

Multimodal, generative, and responsible affective computing aims to enhance people's lives. In recent years, the AI revolution has already begun to impact daily life, with virtual assistants being deployed across various sectors such as healthcare, banking, transportation, and education. It is clear that, in the near future, humans may interact with AI-powered systems as much or maybe even more than direct human-to-human interactions. Affective computing has numerous applications, including innovative approaches to forecasting and preventing anxiety, stress, and mental health issues; enhancing robotic empathy; assisting individuals with communication, behavior, and emotion regulation challenges; and promoting awareness of health and well-being. Many of these applications require enhanced control and protection of sensitive, private, and personal data. Therefore, it is crucial to further develop the creation, evaluation, and deployment of emotionally intelligent systems that are both responsive and responsible. Additionally, improving the accuracy and interpretability of emotion prediction results can significantly enhance the application of this technology in the downstream tasks mentioned above. MRAC'25 is the continuation of MRAC'23 and MRAC'24. Through this workshop, we aim to bring together researchers to discuss the potential and development of affective computing.

Ziyu Wei, Luting Wang 0001, Chen Gao 0005, Hongliang Huang, Jiaqi Liu 0006, Li Wen, Si Liu 0001

Embodied intelligence, evolving from rule-based control to learning-driven systems, has primarily focused on rigid robots, whose limitations in flexibility and adaptability have spurred research into soft-bodied platforms inspired by biological organisms. Soft robots offer adaptive, safe solutions for human collaboration and complex environments but face challenges from underactuation and nonlinear dynamics. This workshop centers on multimodal perception and decision-making in soft robotics, gathering researchers to explore cutting-edge technologies, challenges, and solutions across areas like embodied navigation, manipulation, and control, with presentations and discussions on novel findings and methodologies.

Sherzod Hakimov, David Semedo, Eric Müller-Budack, Marc A. Kastner 0001, Takahiro Komamizu

Multimodal human understanding is an evolving interdisciplinary field integrating computer science, psychology, and social sciences to model human perception, behaviour, and biases in multimodal data. While recent advancements in multimodal learning excel in tasks like image-text synthesis, they often overlook nuanced human-centric dynamics---such as cultural, political, and individual influences on how modalities (e.g., text and images) interact, complement, or contradict each other. The 4th International Workshop on Multimodal Human Understanding (MUWS) aims at addressing these challenges, fostering novel solutions that explicitly model human perception, behaviour, and biases in multimodal data, with a particular emphasis on real-world challenges in web and social media analysis. This year edition covers two tracks: (1) human-centred multimodal understanding, such as quantifying social biases, analysing sentiment and hate speech, and modelling cross-modal interactions through interdisciplinary theories (e.g., semiotics, gestalt psychology); and (2) Multimodal understanding of global events, supported by a newly curated dataset covering news articles with diverse stances, which facilitates research on cultural framing, societal impact, and bias mitigation in vision-language models. The event features two keynotes from renowned experts from journalism and computer science, research presentations for six accepted papers, and interactive discussions to explore and discuss cutting-edge methodologies and applications in multimodal human understanding. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3728481

Taras Kucherenko, Alice Delbosc, Rajmund Nagy, Laura B. Hensel, Youngwoo Yoon, Oya Çeliktutan, Gustav Eje Henter

Imbuing embodied agents with non-verbal behavior offers significant benefits for agent-human interactions. Despite extensive research on the creation of non-verbal behaviors, the field lacks a standardized benchmarking practice. Researchers rarely compare their findings with previous studies, and when they do, the comparisons are often not methodologically aligned. The GENEA Workshop 2025 aims to bring together the non-verbal behavior generation community to discuss major challenges and solutions in the field and determine the most effective ways to advance it. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746268.

Amit Kumar Jaiswal 0001, Thomas Mandl 0001, Gautam Kishore Shahi, Durgesh Nandini, Haiming Liu 0002

With the advancement of digital technologies and gadgets, online content has become easily accessible. At the same time, harmful content also spread widely. There are different harmful content types present on various platforms in multiple languages. The topic of harmful content is broad and covers multiple research directions. Users of platforms are affected by all of them. In research, the different forms are mostly analysed separately, e.g. misinformation, cyber-bullying and hate speech. Most research has been conducted for only one platform, for a monolingual situation or on a particular issue. Counter-measures like blocking are down-ranking can make harmful content spreaders to switch platforms and languages to continuously reach a user base. Harmful content does not only appear on social media but also on news media. Spreader share harmful content in posts, news articles, comments and hyperlinks. There is a great need to study harmful content across platforms, languages, and topics. We plan to bring the research on harmful content under one umbrella such that different approaches and novel methods can be shared. The workshop will also cover the currently ongoing issues of war and elections. We propose the workshop, DHOW: Diffusion of Harmful Content on Online Web, which brings together the research on different topics of harmful content. We expect to discuss innovative research work and future research directions. The proposed workshop is the next iteration of DHOW 2024. https://dhow-workshop.github.io previously organized at ACM WebSci 2024 in Stuttgart, Germany.