论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 9 / 81 页

Tianxing Zhou, Chengkai Xu, Xinyue Yao

So Long is a VR narrative reimagining Voyager 1's final contact. Users decode the Golden Record, embody the probe, and record messages into an evolving interstellar archive. Through AI-generated environments and embodied interaction, the experience reframes space exploration as participatory memory-making. This paper presents its design, technical approach, and user insights within immersive storytelling and digital heritage.

Meichun Cai, Yiou Wang

Mixanthropy is a holographic installation that invites viewers on a shapeshifting journey through human and nonhuman states. This work features digital chimeras that morph in response to the viewer's presence. It explores the body in transition, challenging anthropocentric boundaries and evoking a living, animistic space where form, identity, and consciousness remain in constant flux.

Yifan Chen 0006

Echoes of the Creator is a single-user immersive VR experience that simulates creator-user role shifts through spatial storytelling and multimodal interaction. By fostering embodied empathy and reflecting on power dynamics, it offers insights into co-creative processes. AI-driven companions guide the user and provide narrative feedback throughout the journey. The prototype lays the foundation for future multi-user experiences, inviting new forms of collaborative world-building and creative reflection.

Mingdong Song, Yufei Huang 0022

This project presents an interactive installation that integrates artificial intelligence and visual art, employing immersive environments, AI-driven emotional dialogue systems to explore the possibility of emotional interactions between humans and AI in romantic scenarios. Participants enter a virtual space to engage in romantic conversations with an AI virtual character, sharing their emotions. Meanwhile, external observers use VR devices to experience the interaction from the AI's perspective, enabling a philosophical reflection through a third-person viewpoint. This dual-layered experience ultimately unveils a crucial question: In an era where human-machine interactions blur the boundaries of subjectivity, are emotions the product of genuine connections with an ''other,'' or do they merely reflect a narcissistic self-satisfaction shaped by technology? Through artistic expression, this project provokes profound reflections on subjectivity, the meaning of emotions, and the boundaries of human-AI relationships in a posthumanist context, challenging traditional definitions of emotional essence and offering groundbreaking perspectives.

Xuanyang Huang, Wei Huang

Mirage is a computational film installation that explores how generative AI can fictionalize temporal sequences within global surveillance systems. Constructed from real surveillance frames captured via public webcam infrastructure, Mirage interpolates the time gaps between frames using AI-generated visual and narrative content. Rather than creating fictional worlds from scratch, the work embeds speculative micro-narratives into the neutral flow of surveillance imagery-raising critical questions about visual truth and the narrative potential of machinic vision.

Raphaëlle Lemaire, Azamat Kaibaldiyev, Eléonore Mariette, Débora Viglieri, Alexis Lechervy, Fabrice Maurel, Gaël Dias, Jérémie Pantin, Gaëtane Blaizot, Véronique Agin 等

As Picasso said, a painting lives only through the one who looks at it. To materialize this thought, we propose to automatically produce artworks that visually transform paintings by amplifying and distorting the most observed areas by viewers. Our work is based on a study conducted at the Caen Museum of Fine Arts in France. During the study, 151 participants were equipped with eye-tracking glasses, and observed various paintings, first alone and then in pairs. Based on the fixation and gaze path stored data, we first generate saliency maps that reflect the visual attention given to each painting. These maps are then used to fine-tune the UNETRSal model, a neural network designed to predict saliency maps, in order to align its outputs with human visual patterns observed during the experiment. The saliency maps generated are subsequently used to create deformations of the original painting. This overall process gives rise to a new artwork born from the interaction between human gaze and AI-prediction.

Jinfan Liu, Zhangli Hu, Hanqi Chen, Ye Chen 0006, Bingbing Ni, Shuicheng Yan

We introduce the AR2 O Painter, an interactive intelligent system designed for real-time, highly realistic oil painting creation. This Agent can faithfully reproduce any portrait image, allowing users to visually enjoy the stroke-by-stroke painting process immersively as the artwork is completed within two minutes. It consists of two modules: the Oil Painting Stroke Sequence Planner, which performs multi-level semantic-based brushstroke sequence decomposition on portrait images, mimicking the logic of artist painting, and the Oil Painting Rendering Engine, which receives the brushstroke sequence, models the pigment via fluid dynamics, simulates its interaction with the canvas and brush, and applies a tailored PBR model with microfacet BRDF, Fresnel effects, and stroke-level geometry, enabling perceptually plausible gloss and fine-grained surface relief. To the best of our knowledge, it is the first real-time intelligent painting system to generate realistic oil paintings with high interactivity and artistic fidelity. The demo video is available at https://youtu.be/aN-W06GmnP8.

Masatoshi Hamanaka, Gou Koutaki

This paper describes RoboSax Melody Slot Machine, which automatically plays fingering melodies selected using a dial-like controller with stave notation on a tablet screen. All that are required for the saxophone to output the melody selected by the tablet operator are the blowing and tonguing of the saxophonist. RoboSax Melody Slot Machine makes the task of selecting melodies, which was previously possible only for composers, possible for the general people and enables saxophone playing by those who cannot quickly read the note names on the staff.

Zhucun Xue

This paper presents a doctoral research focusing on integrating Retrieval-Augmented Generation (RAG) into video-related multimodal tasks. Existing RAG studies predominantly target text, images, or tabular data, overlooking the unique value of video as a knowledge carrier. We address this gap by: 1) proposing AdaVideoRAG, a framework that adaptively allocates retrieval strategies based on query complexity for long-video understanding; 2) developing REViG (RAG-Enhanced Video Generation) to optimize prompt engineering via retrieved knowledge for controllable video synthesis; 3) constructing the UltraVideo dataset (UHD-4K/8K resolution, 100+ themes, 10 structured captions per video) and HiVU/HiVG benchmarks to evaluate RAG-driven video tasks. Experiments validate the effectiveness of our methods, and we outline future plans to unify video understanding and generation through Agentic RAG for AGI-oriented research.

João Diogo

Video is an essential part of sports interaction among sportspeople. Athletes benefit from self-visualization to correct movement patterns, coaches leverage video analysis to review past events and spectators engage with video to stay connected with their preferred sports. To meet these distinct user needs, it is crucial to establish a holistic approach that considers the intricacies of human interactions within sports contexts. This PhD research adopts a user-centered approach that iteratively creates, develops and evaluates intelligent video-based systems to support meaningful sports interaction. In line with previous work, this research emphasizes how these systems can facilitate visual analysis and knowledge sharing while addressing interaction challenges posed by deploying computer vision in real-world sports scenarios. The findings from this research contribute to the emerging field of SportsHCI, providing design implications for intelligent video-based systems that enhance learning, game analysis and overall interaction across sportspeople. Additionally, it recognizes the significance of machine learning methods in supporting interpersonal collaboration (e.g., athlete-coach), game understanding and personalization in video-based interactions. Situated at the intersection between HCI, computer vision and SportsHCI, this work aligns with core tenets of research in the field that contribute to enhancing user experiences across multimedia applications.

Yuan-Chun Sun

With the continuous development of networking and computing devices, immersive communication has become increasingly viable, enabling users to explore virtual worlds and interact with other users in 6 Degrees-of-Freedom (6DoF). Immersive communication has great potential not only in professional domains, such as medical diagnostics and distance education but also for leisure activities, such as social networks and new media. This proposal aims to develop an immersive communication system leveraging the cutting-edge dynamic 3D Gaussians Splatting (3DGS). Our objective is to design, implement, and evaluate a highly interactive, adaptive, and efficient immersive communication system that maximizes user experience and system performance while supporting heterogeneous hardware platforms, networks, and applications. We identify and tackle three critical challenges in developing such a system: (i) designing deformable 3DGS avatars to enhance real-time user representation, (ii) developing scalable dynamic 3DGS codecs to optimize data transmission and storage efficiency, and (iii) implementing adaptive streaming algorithms to ensure smooth and responsive user experience across diverse usage scenarios. By solving these challenges, our research aims to push the boundaries of real-time 3D streaming and interaction while redefining the future of virtual world over next-generation networks.

Shiqin Liu, Minjun Zhao, Jiajun Bu

In real-world cooking scenarios, users often need to create personalized recipes based on limited ingredients, dietary goals, and various restrictions, such as the availability of equipment, flavor preferences, and health conditions. Existing recipe generation methods lack the flexibility and adaptability to meet these individual needs. We present LetMeCook, an end-to-end interactive system for personalized recipe generation that leverages multimodal perception, hybrid retrieval, and content generation. Given a photo of the user's refrigerator and a dietary profile, LetMeCook detects available ingredients, retrieves relevant recipe candidates, and generatively refines them based on user requirements, such as ingredient substitution and flavor adjustment. The system provides both textual and visual previews of the adapted recipes, offering a highly interactive and user-centric experience.

Qinglan Wei, Ruiqi Xue, Mingyue Liao, Long Ye

This study presents an intelligent planning system (News Video to Propagation Rules Strategy, NV2PRS). The system is based on event chain modeling to achieve automated generation of event-level communication strategies for video news. It consists of two modules: video style feature extraction and knowledge chain matching. First, a multimodal feature analysis engine is used to obtain text semantics, visual features, and communication features. Then, template-based knowledge chain matching is employed to realize the mapping between events and strategies. To optimize the system's practicality, a hierarchical architecture design is adopted, integrating a feature visualization interface and an end-to-end workflow for strategy generation.

Zhifei Xie, Hu Zongzheng, Guibin Zhang, Jialin Zhang, Yue Liao, Chunyan Miao, Shuicheng Yan

We present Pask, a proactive AI agent that provides real-time, context-aware guidance and knowledge support in audio-centric media environments. Unlike passive assistants that follow the ''you ask, I answer'' model, Pask shifts toward ''answering before asking'' by continuously monitoring live audio, anticipating user needs, and proactively offering conceptual explanations and semantic clarifications. It integrates three core components: a silent copilot for in-situ explanation, a structured knowledge base for factual grounding, and a private memory module for personalized adaptation. Pask enhances comprehension and communication in scenarios such as online learning, media consumption, and live meetings through sustained, intelligent guidance. A live demo is available at https://www.youtube.com/watch?v=ki_CKiV9Oyk.

Jinzhao Zhou, Daniel Leong, Zehong Cao, Thomas Do, Sheng-Fu Liang, Tzyy-Ping Jung, Chin-Teng Lin

This paper presents MindSpeak, a real-time brain-computer interface (BCI) system for recording, processing, and decoding silent speech to enable online multimodal communication between the human brain and a computer, involving both noninvasive multichannel EEG signals and text output. To enable hand-free and brain-only control, our system incorporates steady-state visual evoked potential (SSVEP) for users to select incomplete sentences from a predefined pool and confirm the correctness of decoded words. An intuitive graphical interface is designed for natural communication. We evaluate the effectiveness of our real-time BCI system, which achieves 77.3% accuracy in decoding silent speech and 98.9% accuracy in SSVEP-based selection and confirmation of correct sentences. Unlike existing BCI systems, the presented MindSpeak system significantly expands the application scope of existing BCI systems by enabling users to express complete thoughts through a fully BCI-controlled interactive interface. Our demonstration video is on: https://youtu.be/B1wt1dmCCrg.

Alexander Filonenko, Ilya Makarov, Andrey V. Savchenko

FaceCluster is an interactive photo management system that leverages our enhanced KP-RPE face recognition model with Embedding Statistical Regularization to organize personal photo collections automatically. Unlike existing cloud-based systems that raise privacy concerns, FaceCluster operates entirely locally while demonstrating high performance across multiple challenging benchmarks, including IJB-C (97.25% TAR@0.01%), TinyFace (74.14% Rank-1), and AgeDB (97.78% accuracy) when trained on the WebFace4M dataset. The demo showcases real-time face detection, clustering, and organization capabilities through an intuitive web interface, enabling users to effortlessly manage large photo collections with a single-command Docker deployment.

Milad Ghanbari, Wei Zhou 0021, Cosmin Stejerean, Christian Timmerer, Hadi Amirpour

We present a physics-driven 3D dart-throwing interaction system for Apple Vision Pro (AVP), developed using Unity 6 engine and running in augmented reality (AR) mode on the device. The system utilizes the PolySpatial and Apple's ARKit software development kits (SDKs) to ensure hand input and tracking in order to intuitively spawn, grab, and throw virtual darts similar to real darts. The application benefits from physics simulations alongside the innovative no-controller input system of AVP to manipulate objects realistically in an unbounded spatial volume. By implementing spatial distance measurement, scoring logic, and recording user performance, this project enables user studies on quality of experience in interactive experiences. To evaluate the perceived quality and realism of the interaction, we conducted a subjective study with 10 participants using a structured questionnaire. The study measured various aspects of the user experience, including visual and spatial realism, control fidelity, depth perception, immersiveness, and enjoyment. Results indicate high mean opinion scores (MOS) across key dimensions.

Peng Jin, Yilin Wen 0009, Mingzhe Yu, Yunshan Ma 0002, Rong Zheng, Jintu Fan, Chong Wah Ngo

With the increasing demand for outfit planning in real-world travel scenarios, the need for constructing a travel fashion wardrobe, a series of outfits tailored to a user's personalization and destination-specific context over a short travel period, has grown significantly. However, existing systems or works often focus on isolated factors and rely on retrieval-based methods, with insufficient utilization of generative models, limiting their adaptability to real-world travel scenarios. To address this issue, this study introduces GenWardrobe, a fully generative system for travel fashion wardrobe construction. GenWardrobe consists of three key modules: user query analysis, fashion knowledge retrieval via retrieval-augmented generation and wardrobe image generation. To facilitate users' usage, we encapsulate the solution into an interactive web application. Expert-level evaluation shows that GenWardrobe significantly outperforms traditional systems in both personalization and visual appeal. PowerPoint file and more materials of Genwordrobe can be found on our Github repository: https://github.com/ShanFengShanFeng/GenWardrobe.

Ruifan Ji, Mingyuan Wu, Bo Chen 0025, Michael Zink, Ramesh K. Sitaraman, Jacob Chakareski, Klara Nahrstedt

We present Anywhere Avatar, a telepresence system that enables full-body and facial avatar reconstruction using a smartphone and a laptop. Users record short videos to generate personalized avatars, which are animated in real time during teleconferencing using webcam-based tracking. Built on pre-trained FLAME and SMPL models, the avatars are rendered in high fidelity using Gaussian splatting. The system runs at near real-time with minimal bandwidth, making expressive 3D telepresence accessible without specialized hardware.

Nhu-Binh Nguyen Truc, Nhu-Vinh Hoang, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le

Line sketches serve as the visual DNA of fashion design, forming the essential foundation where concepts take shape, yet today's digital tools often lack the fluidity, personalization, and intelligence needed to truly support this creative process. We present FashSketch, an interactive, multimedia-driven system that reimagines fashion sketching through the lens of generative AI. Designed with a layer-based creative interface, FashSketch empowers designers to ideate, customize, and iterate on sketches seamlessly. By integrating state-of-the-art generative models, sketch-based retrieval, and large language models, the system supports advanced functionalities such as text-to-sketch generation and context-aware sketch recommendation. FashSketch not only enhances the sketching experience but also opens new multimodal pathways for creative expression, making it a powerful co-creative partner in the early stages of fashion design. Video demo is available at https://youtu.be/BX-Edz7Z7ZY.