论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 10 / 81 页

Jungsu Kim, Jungwoo Huh, Yeseung Park, Seongjean Kim, Jeongwook Choi, Sanghoon Lee 0001

In this demonstration, we present Permission to Dance, an end-to-end dance enhancement system designed to capture, enhance, and analyze user's dance performance. Our system consists of a dance capture module, a dance enhancement module, and a dance feedback module. Using the system, users can acquire their dance data in an enhanced version, followed by textual feedback on how to achieve better dance performance. The demonstration video is available at https://youtu.be/lFw7Xic48KU

Michael Francis Perez, Yichi Yang, Yuheng Zha, Enze Ma, Danish Nisar Ahmed Tamboli, Haodi Ma, Reza Shahriari, Vyom Pathak, Dzmitry Kasinets, Rohith Venkatakrishnan 等

Existing video analysis models often lack explainability, perform poorly on long videos, and frequently hallucinate. Commercial solutions are closed-source and costly. We introduce CReLeRI, an open-source system for action detection in untrimmed videos. CReLeRI segments videos using scene and action transitions, detects actions and their arguments and grounds them in 3D space to improve interpretability and reduce hallucinations. The system promotes transparency and trust in AI-driven analysis of complex, real-world videos. A demonstration video is also available.

Ashan Perera, Md Eimran Hossain Eimon, Juan Merlos, Velibor Adzic, Hari Kalva, Borko Furht

As numerous edge devices start implementing intelligent components, the challenges of energy consumption, bandwidth efficiency, and privacy gain significance. One proposed solution relies on the paradigm of split inference, which optimizes the delegation of the computational load between edge and remote devices. We developed and implemented the standard-compliant split inference system with an encoder and decoder capable of real-time streaming and processing. Our system outperforms state-of-the-art video compression implementations by an average of 83% bitrate reduction, while preserving privacy. We demonstrate the system's real-time performance on consumer devices, with interactive visualizations of object detection and segmentation, incorporating real-time metrics. Demo video: https://youtu.be/bmCbUo_ZWWU

Igor Abramov

Existing EEG-driven image reconstruction methods often overlook spatial attention mechanisms, limiting fidelity and semantic coherence. To address this, we propose a dual-conditioning framework that combines EEG embeddings with spatial saliency maps to enhance image generation. Our approach leverages the Adaptive Thinking Mapper (ATM) for EEG feature extraction and fine-tunes Stable Diffusion 2.1 via Low-Rank Adaptation (LoRA) to align neural signals with visual semantics, while a ControlNet branch conditions generation on saliency maps for spatial control. Evaluated on THINGS-EEG, our method achieves a significant improvement in the quality of low- and high-level image features over existing approaches. Simultaneously, strongly aligning with human visual attention. The results demonstrate that attentional priors resolve EEG ambiguities, enabling high-fidelity reconstructions with applications in medical diagnostics and neuroadaptive interfaces, advancing neural decoding through efficient adaptation of pre-trained diffusion models.

Anvar Iskhakov, Viktor Kovalev, Vladislav Naumov, Ilya Makarov

Reliable, real-time detection of sperm-whale clicks is essential yet difficult in noisy ocean audio streams. We present the first browser-based pipeline that combines self-supervised embeddings with a BiLSTM to label clicks at millisecond resolution. The system attains 99% F1 on the Watkins benchmark, reduces false alarms by 60% against the best published baseline, and analyses one-second segments in ~40 ms on a consumer GPU. An interactive UI overlays multi-algorithm detections on waveform and spectrogram views with drag-zoom and live streaming. Code, pretrained weights and the public demo are released to advance bioacoustic event detection.

Nils Riekers, Marten Risius, Tong Chen 0005

Multimodal hate speech detection targets offensive content expressed through combinations of modalities such as text and images, which often evade detection when analyzed separately. We introduce MAXplain, an interactive framework that addresses both issues via a configurable LLM-based multi-agent architecture. Specialized agents handle distinct subtasks and exchange information through structured dialogues, enabling intrinsic explainability and improved accuracy. The web interface supports human-in-the-loop interaction, including real-time adjustment of agent behaviors and evaluation rules. A browser plugin enables direct inspection of online content. While demonstrated for hate speech detection, MAXplain also supports rapid prototyping for other multimodal tasks.

Andreas Babic, Xihui Chen, Djordje Slijepcevic, Adrian Jaques Böck, Matthias Zeppelzauer

We present CounterHelp, a mobile app designed to support adolescents and young people in generating customized counter speech in response to hateful comments on TikTok. CounterHelp allows users to specify counter strategies to combat hate speech. After users share a hateful comment with CounterHelp, it retrieves TikTok metadata to capture its context. Leveraging large language models, CounterHelp generates customised and context-sensitive counter speech in one of four predefined styles, particularly tailored for adolescents. A user experience lab study confirms the effectiveness and usability of CounterHelp and the generated counter speech.

Tan-Hiep To, Duy-Khang Nguyen, Minh-Triet Tran, Trung-Nghia Le

Key Opinion Leaders (KOLs) play a crucial role in modern marketing by shaping consumer perceptions and enhancing brand credibility. However, collaborating with human KOLs often involves high costs and logistical challenges. To address this, we present GenKOL, an interactive system that empowers marketing professionals to efficiently generate high-quality virtual KOL images using generative AI. GenKOL enables users to dynamically compose promotional visuals through an intuitive interface that integrates multiple AI capabilities, including garment generation, makeup transfer, background synthesis, and hair editing. These capabilities are implemented as modular, interchangeable services that can be deployed flexibly on local machines or in the cloud. This modular architecture ensures adaptability across diverse use cases and computational environments. Our system can significantly streamline the production of branded content, lowering costs and accelerating marketing workflows through scalable virtual KOL creation. Video demo is available at https://youtu.be/uXpXmEbjg3M.

Yaojie Li, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, Wu Liu 0005, Ting Yao 0003, Tao Mei 0001

Recent advances in diffusion-based image editing models have demonstrated remarkable success. However, these models primarily rely on high-quality textual prompts to guide image manipulation, creating a significant barrier for non-expert users. In this demonstration, we present an exemplar-based image editing framework named Edit-by-Example, which eliminates the reliance on textual prompts, requires only a single pair of before-and-after images to encapsulate the desired editing effect that can readily be applied on the user-provided query image without any model fine-tuning. Technically, our framework comprises two components: an Adaptive Editing Policy Module (AEPM) and a Generation Module (GM). The AEPM jointly analyzes cross-image relationships in exemplar pairs and query image content to derive optimal editing directions, while GM executes these policies through an off-the-shelf image editor with optional semantic alignment verification. We introduce EEdBench, a comprehensive benchmark for exemplar-based image editing containing 1,500 test cases across 15 categories. Experiments demonstrate that our framework outperforms existing prompt-free methods in editing direction accuracy (S-Visual) and fidelity (FID).

Hadi Amirpour, Doris Putzgruber-Adamitsch, Klaus Schoeffmann

Cataract surgery is the most frequently performed surgical procedure worldwide, involving the replacement of a patient's clouded eye lens with a synthetic intraocular lens to restore visual acuity. Although typically brief, the operation consists of distinct phases that demand precision and extensive training, traditionally constrained by the limitations of real-time observation under a microscope. To enhance learning and procedural accuracy, modern advancements in stereoscopic video capture and head-mounted displays (HMDs) offer a promising solution. This paper demonstrates the application of stereoscopic cataract surgery videos, visualized through Apple Vision Pro (AVP) and Meta Quest 3, to provide immersive 3D perspectives that enhance depth perception and spatial awareness. An expert evaluation study with experienced surgeons indicates that stereoscopic visualization significantly improves comprehension of spatial relationships and procedural maneuvers, suggesting its potential to revolutionize surgical education and real-time guidance in ophthalmic surgery.

Mikhail Mozikov, Daniil Orekhov, Ivan Nasonov, Konstantin Baltsat, Vladislav Pedashenko, Dmitrii Abramov, Nikita Severin, Yury Maximov, Andrey V. Savchenko, Ilya Makarov

This paper presents HL-EAI, a multimodal framework for studying emotion-driven cooperation in human-LLM, human-human, and LLM-LLM interactions via dynamic game-theoretic tasks. HL-EAI integrates emotional prompting, emotion recognition, and expressive AI avatars to enable bidirectional emotion transfer. By modeling affective influence on trust and alignment, it provides a testbed for developing emotionally aware agents. Our demo shows how multimodal emotional cues shape cooperation, advancing socially intelligent, human-compatible AI for interactive multimedia systems.

Zhaofan Qiu, Zijian Gong, Yingwei Pan, Ting Yao 0003, Tao Mei 0001

This paper demonstrates a pioneering unified multimodal agent that transforms complex visual content creation into an intuitive, conversational experience, allowing users to talk, imagine, and evolve their ideas. Overcoming the limitations of fragmented multimodal technique tools, our system seamlessly integrates text-to-image generation, instruction-based image editing, text/image-to-video generation, and interactive understanding within a single AI interface. Users of all skill levels can perform sophisticated visual tasks using natural language and visual inputs. The system's architecture features a central Coordinator module processing multimodal inputs and directing tasks to Generation or Chat pathways. For Generation, a Planner utilizes our state-of-the-art specialized models in image/video generation and image editing, while the Chat function facilitates clarification and collaboration. The interactive demonstration will showcase intuitive multimodal input, seamless real-time content creation/editing, dynamic interactive understanding, and a unified workflow. This agent pioneers a new way for accessible, interactive visual storytelling and collaborative content creation in multimodal generative AI.

Xiao Chen 0022, Wenrui He, Meng Wang 0064, Zhanbin Hu, Chaoquan Shen, Qiang Zhu

In this paper, we present PrivEdit, a zero-shot, interactive image privacy editing system specifically designed for automated sensitive information desensitization. As social networks and smart devices proliferate, the risk of unintended privacy leakage grows, driving demand for personalized, controllable protection tools. PrivEdit is powered by natural-language instructions and integrates a Recognize-Anything model for robust detection and classification of sensitive objects (e.g., faces, license plates, ID cards), followed by GroundingDINO and SAM for high-precision mask extraction. User intents are parsed and disambiguated via GPT-4o, enabling selective target confirmation and iterative refinement. Finally, our editing module performs localized edits-such as adjustable blurring, mosaicking, or replacement via generative editing. With support for multi-round feedback and real-time modification, PrivEdit seamlessly handles both pre-recorded images and live streams, making it ideal for social-media pre-publishing, privacy data desensitization in enterprise or healthcare contexts, and intelligent surveillance applications. By unifying detection, segmentation, intent parsing, and localized editing into one coherent interface, PrivEdit delivers an end-to-end solution for safeguarding visual data. Supplementary materials including the demo video and slides are available at: https://drive.google.com/file/d/13jFBmYgZgxYQLPIAqCeaQhcyzZTHpf7N/view?usp=sharing

Rongyu Zhang, Zhanbin Hu, Jiamu Wang, Qiang Zhu

Visual anomaly detection (VAD) aims to identify image regions that deviate from established normal patterns. Existing methods often rely on domain-specific training and follow a ''one-class-one-model'' paradigm, limiting scalability. We propose Omni-LLaMA-AD, the first unified multimodal large language model for open-set anomaly detection, capable of handling diverse domains with minimal supervision. Built on a pretrained LLaMA backbone, the model uses a VQGAN-based tokenizer and supports joint vision-language generation. Trained via vision-language alignment and instruction tuning, it achieves effective anomaly detection with only a few normal samples and no domain-specific fine-tuning. Our demo showcases the model's ability to generate high-quality anomaly masks across industrial, medical, and logical datasets, highlighting its strong cross-domain generalization and interactive dialogue-based user experience.

Wei Cai, Jian Zhao 0013, Yuchu Jiang, Tianle Zhang, Xuelong Li 0001

Large Vision-Language Models face growing safety challenges with multimodal inputs. This paper introduces the concept of Implicit Reasoning Safety, a vulnerability in LVLMs. Benign combined inputs trigger unsafe LVLM outputs due to flawed or hidden reasoning. To showcase this, we developed Safe Semantics, Unsafe Interpretations, the first dataset for this critical issue. Our demonstrations show that even simple In-Context Learning with SSUI significantly mitigates these implicit multimodal threats, underscoring the urgent need to improve cross-modal implicit reasoning.

Chaolong Yang, Yinuo Guo, Kai Yao, Yuyao Yan, Jie Sun 0024, Kaizhu Huang

This work presents KDTalker++, a real-time system for generating talking portrait videos from a single image using audio or text input. Built on a keypoint-based spatiotemporal diffusion model, it adds voice cloning, background editing, and fine-grained expression control. The demo is available at https://kdtalker.com. A live presentation video is available at https://drive.google.com/file/d/1N4Ggu0Y32DTsV3mbKhXS4kGY6l2ZYKSp/view.

Mohan Zhang, Qianqian Hu, Chuhan Li, Yanxiu Dan, Shenglan Cui, Fang Liu 0002

This paper presents CrePoster, a data-driven framework to generate aesthetic posters for Chinese cultural relics, aiming to enhance the exhibition experience and promote cultural spread. CrePoster comprises three modules: (1) object segmentation module, (2) content generation module, and (3) poster generation module. Upon processing a cultural relic image, the object segmentation module first leverages a cascaded U2Net-SAM structure to obtain the visual target. Secondly, the content generation module utilizes a multi-target learning-enabled caption generator to produce professional captions. Thirdly, the Multimodal Large Language Model (MLLM) based poster generation module adaptively creates aesthetic parameters, including layout and color scheme, ultimately rendering them into refined posters.

Liang Xu, Songkai Jia, Cathal Gurrin, Allie Tran

We present SLIVeR (Someone else's Lifelog in Virtual Reality), an interactive system that reimagines lifelog data through narrative-based Virtual Reality (VR). Instead of passively viewing chronological data, users navigate personal history through cinematic scenes and existential prompts embedded in a memory-reconstruction storyline. Guided by life-oriented questions, users progress from disorientation to recollection using curated lifelog clips. Built in Unity and deployed on Meta Quest 3, SLIVeR transforms lifelogging into a reflective, game-like journey that blends storytelling, gamification, and immersive visuals to enhance user engagement.

Ziqin Wang, Jinyu Chen, Xiangyi Zheng, Qinan Liao, Linjiang Huang, Si Liu 0001

Unmanned Aerial Vehicles, operating in environments with relatively few obstacles, offer high maneuverability and full three-dimensional mobility. This allows them to rapidly approach objects and perform a wide range of tasks often challenging for ground robots, making them ideal for exploration, inspection, aerial imaging, and everyday assistance. In this paper, we introduce AirStar, a UAV-centric embodied platform that turns a UAV into an intelligent aerial assistant: a large language model acts as the cognitive core for environmental understanding, contextual reasoning, and task planning. AirStar accepts natural interaction through voice commands and gestures, removing the need for a remote controller and significantly broadening its user base.It combines geospatial knowledge-driven long-distance navigation with contextual reasoning for fine-grained short-range control, resulting in an efficient and accurate vision-and-language navigation (VLN) capability. Furthermore, the system also offers built-in capabilities such as cross-modal question answering, intelligent filming, and target tracking. With a highly extensible framework, it supports seamless integration of new functionalities, paving the way toward a general-purpose, instruction-driven intelligent UAV agent.The supplementary PPT is available at https://buaa-colalab.github.io/airstar.github.io.

Aishan Liu, Jiakai Wang, Tianyuan Zhang 0004, Hainan Li, Jiangfan Liu 0001, Siyuan Liang 0004, Yilong Ren, Xianglong Liu 0001, Dacheng Tao

Evaluating and ensuring the adversarial robustness of autonomous driving (AD) systems is a critical and unresolved challenge. This paper introduces MetAdv, a novel adversarial testing platform that enables realistic, dynamic, and interactive evaluation by tightly integrating virtual simulation with physical vehicle feedback. At its core, MetAdv establishes a hybrid virtual-physical sandbox, within which we design a three-layer closed-loop testing environment with dynamic adversarial test evolution. This architecture facilitates end-to-end adversarial evaluation, ranging from high-level unified adversarial generation, through mid-level simulation-based interaction, to low-level execution on physical vehicles. Additionally, MetAdv supports a broad spectrum of AD tasks, algorithmic paradigms (e.g., modular deep learning pipelines, end-to-end learning, vision-language models). It supports flexible 3D vehicle modeling and seamless transitions between simulated and physical environments, with built-in compatibility for commercial platforms such as Apollo and Tesla. A key feature of MetAdv is its human-in-the-loop capability: besides flexible environmental configuration for more customized evaluation, it enables real-time capture of physiological signals and behavioral feedback from drivers, offering new insights into human-machine trust under adversarial conditions. We believe MetAdv can offer a scalable and unified framework for adversarial assessment, paving the way for safer AD. Our demo can be found at https://sites.google.com/view/metadv-demo-video.