论文检索

输入标题、作者或关键词,从 1,014 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,014篇论文
第 5 / 51 页

Yang Chen 0048, Jingwen Chen 0001, Yingwei Pan, Xinmei Tian 0001, Tao Mei 0001

We demonstrate an automatic 3D creation system, which can create realistic 3D assets solely from a text or image prompt without requiring any specialized 3D modeling skills. Users can either describe the object they envision in natural language or upload a reference image that records what they have seen with the phone. Our system will generate a high-quality 3D mesh that faithfully matches the users' input. We propose a coarse-to-fine framework to achieve this goal. Specifically, we first obtain a low-resolution mesh instantly by utilizing a pre-trained text/image conditional 3D generative model. Using such coarse mesh as the initialization, we further optimize a high-resolution textured 3D mesh with fine-grained appearance guidance from large-scale 2D diffusion models. Our system can create visually-pleasing results in minutes, which is significantly faster than existing methods. Meanwhile, the system ensures that the resulting 3D assets are precisely aligned with the input text or image prompt. With these advanced capabilities, our demonstration provides a streamlined and intuitive platform for users to incorporate 3D creation into their daily lives.

Qiuyun Zhang, Bin Guo 0001, Lina Yao 0001, Han Wang 0005, Ying Zhang 0047, Zhiwen Yu 0001

Layout generation is important in the field of graphic design and has attracted intensive research attention recently. To further prompt human-computer interactions, we construct the ALDA to assist users throughout the design process, which achieves adaptive and diverse content expansions upon only one element for beginners and generates high-quality posters subsequently. Specifically, to obtain diverse contents, we propose aesthetic-aware design graphs (AGs) for effective poster representation and propose a self-constrained blending strategy upon related examples. In addition, we build a novel layout generator to better arrange elements conditioned on our AGs. Finally, we implement ALDA as an online tool with a set of controllable factors to enhance its practicality.

Ming Feng, Kele Xu, Hengxing Cai

Sound event detection (SED) refers to recognizing the sound events in a continuous audio signal, which has drawn increasing interest during recent decades. The applications of SED seem to be evident in many fields, ranging from surveillance to monitoring applications. Despite the sustainable efforts that have been made, most of the previous attempts are performed on the closed-set, as only fixed and known sound event classes can be employed during the training. In this paper, we present our incremental few-shot SED framework under the open-set settings, as a practical machine listening system should be able to address unknown sound events. Specifically, an explicit learning and calibration-based multi-stage learning framework is utilized to address the challenges of catastrophic forgetting, and aim to achieve a better trade-off between stability and plasticity. To compress the model efficiently, the model prune and self-distillation paradigm are combined used for the model compression, thus our system can be deployed for the resource-limited devices. Our framework can also provide an uncertain estimation for the inference. Lastly, an interactive interface is presented to demonstrate the functions of our system.

Yuya Moroto, Rintaro Yanagi, Naoki Ogawa, Kyohei Kamikawa, Keigo Sakurai, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama

Multimedia content recommendation needs to consider users' preferences for each content. Conventional recommender systems consider them with wearable sensors, however, wearing such sensors can lead to a burden on users. In this paper, we construct a recommender system that can explicitly estimate users' preferences without wearable sensors. Specifically, by constructing lightweight but strong machine learning models suitable for our system, the users' interest levels for contents can be estimated from facial images obtained from a widely used webcam. In addition, through the interaction that the user selects displayed contents, our system finds the tendency of personal preferences for recommending contents with high user satisfaction. Our system is available on https://www.lmd-demo.org/2022/start_eng.html.

Dongkai Wang, Shiliang Zhang, Yaowei Wang 0001, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001

Human-centric visual analysis is a fundamental task for many multimedia and computer vision applications, such as self-driving, multimedia retrieval, and augmented reality, etc. Based on our recent research efforts on fine-grained human visual analysis, we develop a robust and efficient human-centric visual analysis system named as HumVis. HumVis is built on a simple yet efficient contextual instance decoupling (CID) module, which can effectively separate different persons in an input image and output corresponding person structure information for visual analysis. Based on CID, HumVis achieves accurate multi-person pose estimation, multi-person foreground segmentation, multi-person part segmentation and 3D human mesh recovery for user-uploaded images/videos and support live stream presentation.

Zeyu Jin, Zixuan Wang 0026, Qixin Wang 0002, Jia Jia 0001, Ye Bai 0001, Yi Zhao 0006, Hao Li 0078, Xiaorui Wang

Lyrics and music are both significant for a singer to perform a song. Therefore, it is important in singer's motion generation to model both semantic and acoustic correlation with motions at the same time. In this paper, we propose HoloSinger, a novel comprehensive system that synthesizes singing motions according to the given song. Additionally, we present singing avatar with octahedral holographic projection. For singing motion generation, we introduce a Transformer-VAE generative model to decompose lyrics and music, then fuse their impacts to synthesize singer's motions. Extensive experiments and user studies show that our method automatically generates realistic motions that adhere to musical choreography and reflect the lyric semantics appropriately. Furthermore, we design a desktop-level holographic projection device with an octahedral structure. It achieves high-definition holographic projection effects with smaller volume, larger imaging area ratio, and the ability of real-time AI interaction.

Zhanbin Hu, Jianwu Wu, Danyang Gao, Yixu Zhou, Qiang Zhu

The traditional controllable face generation refers to the controllability of coarse-grained ranges such as facial features, expression postures, or viewing angles, but specific application scenarios require finer-grained control. This paper proposes a fine-grained and controllable face generation technology, CFTF. CFTF allows users to participate deeply in the face generation process through multiple rounds of language feedback. It not only enables control over coarse-grained features such as gender and viewing angle, but also provides flexible control over details such as hair color, accessories, and iris color. We apply CFTF to the suspect portrait scene, and perform multiple rounds of human-computer interaction based on the eyewitness's painting sketch of the suspect and descriptions of their facial features, realizing the "Human-in-the-loop" collaborative portrait drawing.

Yuki Konishi, Panote Siriaraya, Da Li 0008, Katsumi Tanaka, Yukiko Kawai, Shinsuke Nakajima

In recent years, the number of people who run for the purpose of improving their health and physical fitness has been increasing, but it is not easy to continue running. Therefore, we believe that it is very important to develop a running support system. In our previous study, we developed a running support system and an application that enables users to run with a virtual runner created by using their past running data in an acoustic augmented reality space. However, this system has the problem that it can only race against the past running records of oneself or one's acquaintance, and can only race against a single running record. Therefore, we considered it necessary to develop a system that allows users to race against an unspecified number of users and that allows users to arbitrarily select the distance and number of times they wish to race. In this paper, we propose an online marathon system that enables large-scale online races and examine the effect of the number of competitors on user motivation.

Hao Wu 0089, Yueyao Li, Yan Zhuang 0006, Xinyao Sun, Wei Cai 0002

Mainstream blockchain games have drawn criticism for prioritizing economic systems over gameplay experience. Influenced by these economically-centered games, existing research on blockchain games predominantly focuses on the financial sector. We have developed BranchClash, a fully on-chain tower defense game on the Sepolia testnet of Ethereum. It introduces chain collaboration, a novel non-economically-centered game mechanism inspired by blockchain technology. BranchClash aims to expand unique game mechanics in blockchain games and explore innovative cooperative modes within the decentralized ecosystem.

Yu-Hsuan Chen, Chen-Wei Fu, Wei-Lun Huang, Ming-Cong Su, Hsin-Yu Huang, Andrew Chen, Tse-Yu Pan

Volleyball, a sport characterized by unpredictable factors such as ball trajectory, teammate actions, and strategic positioning, presents a challenge when it comes to modeling and training due to its high levels of complexity. Successful gameplay relies on the coordinated efforts of all team members in the receiving, setting, and attacking phases. In real-life competitions, the setter's on-ball ability and decision-making are particularly crucial to the team's offensive success: To improve the training of setters in observing player movements while running and making informed attacking decisions, we propose the design of a virtual reality (VR) system which aims to enhance players' setting skills and strategic thinking to achieve more successful offensive plays with a lower cost.

Mizuki Takenawa, Naoki Sugimoto, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa

We propose a system to generate 360° realistic virtual worlds (360RVW) for the interactive spatial exploration of omnidirectional street-view videos. Our 360RVW enables users to explore photorealistic scenes with digital avatars, and interact with others. To create the virtual worlds our system only requires 360° videos with annotations of the start and end camera coordinate as input. We first detect street intersections to divide the input videos and remove the camera operator from the recordings using a video completion technique. Next, we analyze the 3D structure of the scene using semantic segmentation to define walkable areas. Finally, we render the environment using an ellipsoid projection surface to achieve a more realistic integration of the avatar into real-world 360° videos. The whole process is largely automated, enabling users to produce realistic and interactive virtual worlds without specialized skills or time-consuming manual interventions.

Yi Han, Kaidong Li, Zihan Song 0003, Wei Feng, Xiang Cao, Shida Guo, Xin Wang 0019, Xuguang Duan, Wenwu Zhu 0001

We present H2V4Sports, a real-time horizontal-to-vertical video converter specifically designed for sports live broadcasts. With the increasing demand of smartphone users who prefer to watch sports events on their vertical screens anywhere, anytime, our platform provides a seamless viewing experience. We achieve this by fine-tuning and pruning an object detector and tracker, which enables us to provide real-time, accurate key-object tracking results despite the complexity of sports scenes. Additionally, we propose a video virtual director platform that captures the most informative vertical zones from horizontal video live frames using various director logic for a smooth frame-to-frame transition. We have successfully demonstrated our platform in two popular sports: basketball and diving, and the results indicate that our technology delivers high-quality vertical scenes that are beneficial for smartphone users and other vertical scenarios.

Djamahl Etchegaray, Yadan Luo, Zachary FitzChance, Anthony Southon, Jinjiang Zhong

Road surveying plays a vital role in effective road network management for local governments. However, current practices pose challenges due to their costly, time-consuming, and inaccurate nature. In this paper, we propose an automated survey platform that supports weed, defect and asset monitoring with instance segmentation models. Empowered by recent advancements in vision-language models (VLMs), our solution offers improved flexibility for novel tasks with a limited label set. For domain specific classes, such as pavement cracks and potholes, we train a detector to identify their location given our sparsely annotated images, and alleviate false-positives by rejecting predictions outside regions of interest identified by VLMs. The proposed system directly involves managers in the survey process through a mobile application. The application allows three core functions: 1) capture and cloud upload, 2) real-time survey trajectory monitoring, 3) open-vocabulary detection.

Junchen Zhu, Huan Yang 0005, Wenjing Wang 0001, Huiguo He, Zixi Tuo, Yongsheng Yu, Wen-Huang Cheng, Lianli Gao, Jingkuan Song, Jianlong Fu 等

Videos for mobile devices become the most popular access to share and acquire information recently. For the convenience of users' creation, in this paper, we present a system, namely MobileVidFactory, to automatically generate vertical mobile videos where users only need to give simple texts mainly. Our system consists of two parts: basic and customized generation. In the basic generation, we utilize the pretrained image diffusion model, and adapt it to a high-quality open-domain vertical video generator. As for the audio, by retrieving from our big database, our system matches a suitable background sound for the video. Additionally to produce customized content, our system allows users to add specified screen texts for enriching visual expression, and specify texts for automatic reading with optional voices as they like.

Zheng Zhang 0006, Songling Chen, Mixiao Hou, Guangming Lu 0002

In this paper, we present a multimodal emotion analysis platform, which can flexibly capture, detect and analyze the emotions of video object with multiple modalities under different situations, including offline and online application scenarios. This system can visualize the dynamic effects of different types of emotions from both multimodal and unimodal circumstances. The presented emotion analysis results show instant and time series states in both specific modality and multiple modalities. Our system fills the current research and application gaps in multimodal emotion analysis with an interactive interface. Notably, the constructed system can adaptively process pre-recorded video clips as well as collected real-world data with excellent practicality and interactivity.

Qinghao Ye, Haiyang Xu 0001, Ming Yan 0008, Chenlin Zhao, Junyang Wang 0001, Xiaoshan Yang, Ji Zhang 0011, Fei Huang 0002, Jitao Sang 0001, Changsheng Xu

Inspired by the recent developments of large language models (LLMs), we propose mPLUG-Octopus, a versatile conversational assistant designed to provide users with coherent, engaging, and helpful interaction experiences in both text-only and multi-modal scenarios. Unlike traditional pipeline chatting systems, mPLUG-Octopus offers a diverse range of creative capabilities including open-domain QA, multi-turn chatting, and multi-modal creation, all built with a unified multimodal LLM without relying on any external API. With the modularized end-to-end multimodal LLM technology, mPLUG-Octopus efficiently facilitates engaging and open-domain conversation experience. It exhibits a wide range of uni/multi-modal elemental capabilities, enabling it to seamlessly communicate with users on open-domain topics and engage in multi-turn conversations. It also assists users in accomplishing various content creation and application tasks. Our conversational assistant can also be deployed on smart hardware to drive advanced AIGC applications.

Sandipan Sarma

Human beings possess the remarkable ability to recognize unseen concepts by integrating their visual perception of known concepts with some high-level descriptions. However, the best-performing deep learning frameworks today are supervised learners that struggle to recognize concepts without training on their labeled visual samples. Zero-shot learning (ZSL) has recently emerged as a solution that mimics humans and leverages multimodal information to transfer knowledge from seen to unseen concepts. This study aims to emphasize the practicality of ZSL, unlocking its potential across four different applications in computer vision, namely -- object recognition, object detection, action recognition, and human-object interaction detection. Several task-specific challenges are identified and addressed in the presented research hypotheses. Zero-shot frameworks are proposed to attain state-of-the-art performance, elucidating some future research directions as well.

Yuchen Yang

Situated in the intersection of audiovisual archives, computational methods, and immersive interactions, this work probes the increasingly important accessibility issues from a two-fold approach. Firstly, the work proposes an ontological data model to handle complex descriptors (metadata, feature vectors, etc.) with regard to user interactions. Secondly, this work examines text-to-video retrieval from an implementation perspective by proposing a classifier-enhanced workflow to deal with complex and hybrid queries and a training data augmentation workflow to improve performance. This work serves as the foundation for experimenting with novel public-facing access models to large audiovisual archives.

Ying Fang

The emerging haptic technology has introduced new media perceptions and also increased the immersive experiences of end-users. To date, novel haptic-audio-visual environments and their Quality of Experience (QoE) assessments are still challenging issues. In this work, we investigate the haptic-visual interaction QoE in virtual as well as real-world environments. First, we establish a haptic-visual interaction platform based on a balance ball Virtual Reality (VR) game scene and a haptic-visual interaction platform with data-glove-assisted remote control. Second, we conduct subjective tests to qualitatively and quantitatively analyze the impacts of system-related, user-related and task-related factors on QoE evaluation. Third, we propose learning-based QoE models to effectively evaluate the user-perceived QoE in haptic-visual interaction. In the future work, we aim to focus on the improvement of the two established platforms, with the addition of audio-related influencing factors and more haptic feedback, and optimizing the proposed QoE model for further improvement of haptic-audio-visual interaction applications.

Keke Zhang

Image enhancement techniques play an important role in improving visual quality of media content. Image Quality Assessment (IQA) metrics are essential in comparing and improving image enhancement algorithms. Although these enhancement methods belong to different types of tasks in computer vision, the IQA for these methods has a common problem, namely, their reference information is limited. This poses a challenge to IQA: How to utilize the limited reference information to evaluate the qualities of distorted images. I am dedicated to approaching this challenge throughout my PhD. In this paper, first, I illustrate the motivation of my PhD project. Then, I formulate the problem and the ultimate goal of my PhD research. Next, I introduce some developed results in my research. Finally, I present my ongoing work and future directions.