论文检索

输入标题、作者或关键词,从 1,237 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,237篇论文
第 3 / 62 页

Kaicheng Yu, Zhuang Shao, Siyuan Qi, Dongfang Liu

The tutorial "Large Vision-Language Model in the Society" aims to provide a comprehensive overview of state-of-the-art techniques and applications of large vision-language models (LVLMs), which integrate visual and textual data to transform multimedia research and applications. LVLMs are poised to revolutionize domains such as content creation, social media analysis, education, healthcare, and entertainment by enabling sophisticated content analysis, retrieval, and generation. This tutorial will cover the fundamentals of vision-language integration, state-of-the-art models, training techniques, applications, ethical considerations, and future directions. It is designed to be educational and instructive, providing an in-depth introduction rather than a cursory survey. Attendees will gain practical skills, and insights into the latest research, and engage in interactive sessions to reinforce learning. By addressing both technical and societal aspects, the tutorial will significantly benefit the multimedia community, driving innovation and progress in the field.

Xin Wang 0019, Yuwei Zhou, Hong Chen 0011, Wenwu Zhu 0001

This tutorial focuses on curriculum learning (CL), an important topic in machine learning, which gains an increasing amount of attention in the research community. CL is a learning paradigm that enables machines to learn from easy data to hard data, imitating the meaningful procedure of human learning with curricula. As an easy-to-use plug-in, CL has demonstrated its power in improving the generalization capacity and convergence rate of various models in a wide range of scenarios such as computer vision, natural language processing, reinforcement learning, etc. In particular, CL can also play an important role in multimedia applications. Therefore, it is essential to introduce CL to more scholars and researchers in the machine learning and multimedia community. However, there have been no tutorials on CL for multimedia so far, motivating the organization of this tutorial at ACM Multimedia 2024. To give a comprehensive tutorial on CL for multimedia, we plan to organize it from the following aspects: (1) theories, (2) approaches, (3) applications, (4) tools, and (5) future directions. First, we introduce the motivations, theories, and insights behind CL. Second, we advocate novel, high-quality approaches, as well as innovative solutions to the challenging problems in CL. Then we present the applications of CL in various scenarios, especially multimedia, followed by some relevant tools. In the end, we discuss open questions and future directions in the era of large language models. We believe this topic is at the core of the scope of ACM Multimedia and is attractive to the audience interested in machine learning and multimedia from both academia and industry.

Soyeon Caren Han, Feiqi Cao, Josiah Poon, Roberto Navigli

This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the foundational concepts of multimodality, the evolution of multimodal research, and the key technical challenges addressed by these models. We will cover the latest multimodal datasets and pretrained models, including those beyond vision and language. Additionally, the tutorial will delve into the intricacies of multimodal large models and instruction tuning strategies to optimise performance for specific tasks. Hands-on laboratories will offer practical experience with state-of-the-art multimodal models, demonstrating real-world applications like visual storytelling and visual question answering. This tutorial aims to equip researchers, practitioners, and newcomers with the knowledge and skills to leverage multimodal AI. ACM Multimedia 2024 is the ideal venue for this tutorial, aligning perfectly with our goal of understanding multimodal pretrained and large language models, and their tuning mechanisms.

Wei Gao 0003, Ge Li 0002

Point clouds have the strong capability for modeling 3D objects and scenes, which can be widely used in diverse applications and thus generate the burdens of transmission and storage. Efficient compression algorithms have been explored extensively, and research efforts have also been invested to enhancement algorithms. Moreover, the quality of point clouds can influence 3D analysis tasks, e.g., classification, segmentation, detection, and multimodal understanding, etc. Recent 3D multimodal large models can bring better perception optimizations. This tutorial will provide the fundamental knowledge for point cloud compression, enhancement and applications, and place emphasis on the influences of point cloud quality to human and machine perceptions. We will also discuss the progress of international standards and open source projects for point cloud technologies. From this tutorial, audiences are expected to grasp the basic knowledge and recent progress of point cloud technologies, and promote the research developments in both academia and industrial communities.

Hao Fei 0001, Xiangtai Li, Haotian Liu, Fuxiao Liu, Zhuosheng Zhang 0001, Hanwang Zhang, Shuicheng Yan

Artificial intelligence (AI) encompasses knowledge acquisition and real-world grounding across various modalities, including language, visual, auditory, and sensory data. Multimodal large language models (MLLMs) have thus recently garnered growing interest in both academia and industry, showing an unprecedented trend to achieve human-level AI. This tutorial aims to deliver a comprehensive review of cutting-edge research in MLLMs, focusing on three key areas: MLLM architecture design, instructional learning, and multimodal reasoning of MLLMs. We will explore technical advancements, synthesize key challenges, and discuss potential avenues for future research. All the resources and materials will be made available online. https://mllm2024.github.io/ACM-MM2024

Niccolò Biondi, Simone Ricci, Federico Pernici, Alberto Del Bimbo

In today's multimedia-rich environment, the rapid growth of data poses significant challenges for developing efficient multi-modal retrieval systems essential for retrieving text, images, audio, and video. As data expands, newer, scalable, and high-performance retrieval systems are increasingly necessary. Embedding-based deep neural networks (DNNs) have become key solutions, transforming high-dimensional data into lower-dimensional embeddings for easy comparison and retrieval. However, updating DNNs changes the internal feature representations, necessitating the extraction of new feature vectors for all gallery data, which is costly, especially with gallery sets comprising billions of data. Learning backward-compatible representations addresses this by allowing new representation to be matched with old gallery data without recalculating features. This tutorial aims to equip participants with the knowledge and tools to apply backward-compatible representations, enhancing multimedia retrieval systems' efficiency and scalability. Participants will learn the importance of compatible representations, basic methods and techniques, and explore challenging open questions that are becoming increasingly relevant to multimedia and cross-modal retrieval.

Rahel Arnold, Werner Bailer, Ralph Gasser, Björn Þór Jónsson 0001, Omar Shahbaz Khan, Heiko Schuldt, Florian Spiess 0001, Lucia Vadicamo

The way we create, consume and interact with multimedia content has changed significantly in recent years with the advent of affordable recording devices and easy sharing and access in the form of mobile phones. With the imminent wave of affordable devices that enable mixed reality experiences and the large variety of devices on the market, interaction with multimedia content is expected to continue to evolve rapidly. This will also drastically affect the entire area of multimedia information retrieval in eXtended Reality (XR), for instance by novel ways to express user needs in VR, result presentation that takes the specific capabilities of XR devices into account, and/or result feedback. This tutorial on Multimedia Retrieval in XR discusses and demonstrates existing solutions and highlights key challenges in this evolving field.

Shengzhou Yi, Junichiro Matsugami, Takuya Yamamoto, Toshihiko Yamasaki

Presentation skills, which involve the effective use of verbal and nonverbacl cues, enable audiences to better understand the content being presented. We develope a deep learning-based online assessment system that can objectively evaluate speakers' oral presentations and slide design, providing comprehensive feedback to support their self-practice. For the speaking skill assessment, we construct a multimodal neural network, including LSTMs and attention networks, to analyze the linguistic and acoustic features of oral presentations. The proposed model can predict 14 distinct types of audience impressions with an average accuracy of 85.0%. For the slide design assessment, we propose a method that can analyze slide design based on their visual and structural features, independent of file formats. It can determine whether the slides meet 10 assessment criteria with an average accuracy of 81.7%.

Yuning Wu 0001, Jiatong Shi, Yifeng Yu, Yuxun Tang, Tao Qian, Yueqian Lin, Jionghao Han, Xinyi Bai, Shinji Watanabe 0001, Qin Jin

This research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both continuous and discrete approaches. Specifically, we explore discrete representations derived from SSL models and audio codecs and offer significant advantages in versatility and intelligence, supporting multi-format inputs and adaptable data processing workflows for various SVS models. The toolkit features automatic music score error detection and correction, as well as a perception auto-evaluation module to imitate human subjective evaluating scores. Muskits-ESPnet is available at https://github.com/espnet/espnet.

Mingyuan Wu, Ruifan Ji, Haozhen Zheng, Jiaxi Li, Beitong Tian, Bo Chen 0025, Ruixiao Zhang, Jacob Chakareski, Michael Zink, Ramesh K. Sitaraman 等

We propose an interactive and intelligent hybrid teleconferencing system compatible with Virtual Reality devices. Our system understands meeting contexts and leverages user interactions to enhance better system configuration. Employing interactive scene graphs [11], the system extracts and transmits essential meeting context to users while relaying user interactions back to the streaming systems for user-involved adaptive streaming and foveated rendering. We demonstrate the system's real-time performance and compatibility with commercial VR devices such as the Meta Quest 3.

Liangyu Wang, Yoko Yamakata, Ryoma Maeda, Kiyoharu Aizawa

We developed an application that can easily calculate the nutritional content of a meal by utilizing our multimedia recipe dataset tied to the Nutrition Facts table and an ingredient estimation model. A CLIP-based image recognition model and an ingredient co-occurrence graph assist a user in selecting the appropriate ingredient from over 2,500 ingredients. Unlike traditional food applications, ours calculates nutrition from a list of ingredients, allowing the user to see how each ingredient affects the overall nutrition. The user adjusts the amount of each ingredient and knows how to change their meal to meet their body's needs.

Ying Ma, Xinyan Yang, Aiqi Wang, Jianglin Zeng, Shaofei Liu

In this work, we present a video editing chatbot (VEC) that performs intelligent multimedia editing through natural language dialogue. VEC comprises three modules: instruction analysis, multimedia resources retrieval, and multimedia resources editing. It analyzes user instructions to retrieve relevant multimedia resources from the multimedia database (MMDB), and then applies appropriate editing methods from the multimedia toolbase (MMTB) automatically. To enhance user experience and simplify operation, VEC uses a multi-turn dialogue mechanism to handle complex editing tasks.

Seongjean Kim, Jungwoo Huh, Yeseung Park, Jungsu Kim, Sanghoon Lee 0001

In this demonstration, we present DanceMimic, a real-time dance imitation capture and assessment system designed to enhance the accessibility and learning experience of dancing. Guided by an interactive user interface, novice dancers can simultaneously observe and imitate a selected choreography while listening to the corresponding music. The choreography is captured and compared with the reference for quantitative evaluation of the performance proficiency. Finally, our system retargets the performed dance to a rigged 3D character to provide immersive imitation experience. Demo video is on: https://youtu.be/7nL9YPPRj-4

Xin Jin 0015, Liaoruxing Zhang, Longteng Jiang, Dandan Li

This paper introduces a novel method for enhancing image composition guidance in photography. It utilizes advanced composition rules to guide a Real-Time Detection Transformer (RT-DETR) model in predicting aesthetically pleasing compositions for photographs. Unlike traditional methods constrained by original image boundaries, our approach allows the predicted framing to extend beyond these limits, offering dynamic, real-time guidance for image composition in photography. The system integrates multi-label composition classification and compositional element annotation, using YOLOv8 for key object detection and an enhanced Deep Hough Transform for compositional lines to guide photographers. It provides photographers with real-time guidance for optimal camera adjustments, transforming traditional post-processing tasks into an intuitive, interactive process. This method significantly enhances photographers' flexibility and effectiveness in capturing visually superior photographs.

Zhanbin Hu, Xiaodong He, Renzhou Pan, Xianzhou Zeng, Chenming Fan, Qiang Zhu

In the domain of video generation, Text-to-video suffers from a notable application gap due to lack of audio that harmonizes with the visual content. Current solutions typically dubbing based solely on the original text used for generate video, which causes a mismatch between the video content and audio details, primarily stems from the lack of understanding of the video's visual modality. Leveraging advancements in multimodal large language model and LLM-based Agent, we propose MAF-ID, a multi-agent interactive framework for video dubbing based on deep video understanding. MAF-ID achieves agent collaboration through the autonomous interaction of three agents, to capture a deep understanding of the video visual content from macro to micro, progressively generate sound effects, voice-overs, and background music that is adaptive to the video. By deeply aligning text, video, and audio modalities, our method significantly enhances the fine-grained coordination between video and audio, making it widely available for AI-generated videos, VLOGs, and other video production scenarios requiring dubbing.

Feilin Han, Leping Zhang, Xin Wang 0019, Ke-Ao Zhao, Ying Zhong, Ziyi Su, Tongtong Feng, Wenwu Zhu 0001

Unmanned Aerial Vehicles (UAVs) are necessary across diverse domains, including disaster surveillance and wildlife conservation. However, the development and evaluation of UAV-related algorithms often encounter a significant hurdle: the scarcity of authentic training data. In this paper, we introduce U2USim, a telepresence simulation platform with a dynamic environment, serving as a realistic synthetic data generation, performance evaluation, and visualization tool for UAV-to-UAV (U2U) cooperative learning. This paper presents the architecture, features, and capabilities of U2USim. Leveraging Unreal Engine (UE), AirSim APIs, and ROS (Robot Operating System), our platform enables realistic simulations, mirroring real-world conditions and facilitating research in UAV technology.

Difei Gao, Siyuan Hu, Zechen Bai, Qinghong Lin, Mike Zheng Shou

Graphical User Interface (GUI) Automation has shown significant potential recently. Previous works built GUI Agent systems to handle short-procedure tasks such as element grounding or functional assistance. In this paper, we propose a novel PC-Copilot, AssistEditor, that focuses on automating the video editing workflow. Unlike previous approaches, our system does not require users to input specific commands to control the computer. Instead, users simply describe their requirements, such as the content and style of the video, and upload the necessary materials. The system then autonomously translates these requirements into detailed actions for controlling video understanding models and professional video editing software, e.g., Premiere Pro to produce the final video. This functionality is enabled by a collaborative AI agent framework of multiple GUI agents, each capable of dialogue, knowledge retrieval, and software usage. These agents have distinct roles, including interacting with users to gather requirements, generating storyboards, and performing editing tasks. This approach significantly streamlines the video editing process, making advanced editing accessible to users with varying levels of expertise.

Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Jeonghwan Kim, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang 等

We present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems.

Hang Yuan, Wei Gao 0003, Wenxu Gao

Subjective experiments driven by human visual perception for images aid in the development of related technologies such as analysis, compression, and transmission, which garners substantial attention and research interest. However, the general subjective experiment platform is relatively lacking, and verifying the reliability of annotation data is often difficult. In response to these challenges, an open source subjective experiment platform, namely OpenSEP, is proposed in this paper. Specifically, OpenSEP mainly includes a contrast mode sub-platform that displays dual stimuli, allowing for the simultaneous display of both source stimulus and distorted stimulus for subjective testing. Moreover, a scoring mode sub-platform that displays single stimulus is also provided in OpenSEP. In this mode, subjects can only score the distorted stimulus individually after sequentially viewing the source stimulus and all the distorted stimuli. Besides, OpenSEP constructs a cross-validation sub-platform integrated mainstream Just Noticeable Distortion (JND) algorithms. Within this sub-platform, the reliability of subjective annotation data can be verified based on existing JND algorithms, and the effectiveness of newly proposed modeling algorithms can also be validated. The open source library for OpenSEP is available at https://openi.pcl.ac.cn/OpenDatasets/OpenSEP

Feng Ye, Li Zhang 0006, Chuanmin Jia

Deep video compression has attracted increasing attention in recent years due to its end-to-end optimization ability. However, most existing neural video compression (NVC) models focus on incorporating sophisticated motion or residual coding networks for successive frames leveraging spatial-temporal redundancy removal, neglecting the efficient motion representation and essential structure of scaled prediction for motion dynamics. To resolve this problem, this paper proposed a novel model, named scaled hierarchical bi-directional prediction structure, which effectively captures temporal correlation among frames considering the quality variation when managing the reference frames. This paper first introduces parameter-shared motion codecs and efficient information fusion strategies to obtain predictive features more precisely. Subsequently, scaled motions from temporal contexts are learned as bi-directional prior for motion representation. Additionally, the concept of trustworthy motion modeling is proposed to represent the effectiveness of reference information, measuring the reliability of predictive accuracy in complex motions, camera rotations and occlusions. Extensive experimental results demonstrate that our approach offers significant advantages over state-of-the-art bi-directional NVC models in coding efficiency. The proposed method has been adopted as the latest reference model by Moving Picture, Audio, and Data Coding by Artificial Intelligence (MPAI) end-to-end video coding (EEV) standard. The code is available at: https://github.com/yefeng00/DVC_with_Scaled_Hierarchical_Bi_directional_Motion_Model.