Recent advancements in large-scale video-language models have shown significant potential for real-time planning and detailed interactions. However, their high computational demands and the scarcity of annotated datasets limit their practicality for academic researchers. In this work, we introduce VideoLLaMB, a novel and efficient framework for long video understanding that leverages recurrent memory bridges and temporal memory tokens to enable seamless encoding of entire video sequences with preserved semantic continuity. Central to our approach is a SceneTiling algorithm that segments videos into coherent semantic units, facilitating robust understanding across tasks without requiring additional training. VideoLLaMB achieves state-of-the-art performance, surpassing existing models by 4.2 points on four VideoQA benchmarks and by 2.06 points on egocentric planning tasks. Notably, it maintains strong performance under extreme video length scaling (up to 8x) and excels at fine-grained frame retrieval on our proposed Needle in a Video Haystack (NIAVH) benchmark. With linear GPU memory scaling, VideoLLaMB processes up to 320 frames using a single Nvidia A100 GPU, despite being trained on only 16 frames--offering an unprecedented balance of accuracy, scalability, and cost-effectiveness. This makes it highly accessible and practical for the academic community.
论文检索
输入标题、作者或关键词,从 5,999 篇学术成果中精准定位
One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt Evolution
PDF ↗Polyp segmentation is vital for early colorectal cancer detection, yet traditional fully supervised methods struggle with morphological variability and domain shifts, requiring frequent retraining. Additionally, reliance on large-scale annotations is a major bottleneck due to the time-consuming and error-prone nature of polyp boundary labeling. Recently, vision foundation models like Segment Anything Model (SAM) have demonstrated strong generalizability and fine-grained boundary detection with sparse prompts, effectively addressing key polyp segmentation challenges. However, SAM's prompt-dependent nature limits automation in medical applications, since manually inputting prompts for each image is labor-intensive and time-consuming. We propose OP-SAM, a One-shot Polyp segmentation framework based on SAM that automatically generates prompts from a single annotated image, ensuring accurate and generalizable segmentation without additional annotation burdens. Our method introduces Correlation-based Prior Generation (CPG) for semantic label transfer and Scale-cascaded Prior Fusion (SPF) to adapt to polyp size variations as well as filter out noisy transfers. Instead of dumping all prompts at once, we devise Euclidean Prompt Evolution (EPE) for iterative prompt refinement, progressively enhancing segmentation quality. Extensive evaluations across five datasets validate OP-SAM's effectiveness. Notably, on Kvasir, it achieves 76.93% IoU, surpassing the state-of-the-art by 11.44%.
ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
PDF ↗Perceiving and autonomously navigating through work zones is a challenging and under-explored problem. Open datasets for this long-tailed scenario are scarce. We propose the ROADWork dataset to learn to recognize, observe, analyze, and drive through work zones. State-of-the-art foundation models fail when applied to work zones. Fine-tuning models on our dataset significantly improves perception and navigation in work zones. With ROADWork, we discover new work zone images with higher precision (+32.5%) at a much higher rate (12.8x) around the world. Open-vocabulary methods fail too, whereas fine-tuned detectors improve performance (+32.2 AP).Vision-Language Models (VLMs) struggle to describe work zones, but fine-tuning substantially improves performance (+36.7 SPICE). Beyond fine-tuning, we show the value of simple techniques. Video label propagation provides additional gains (+2.6 AP) for instance segmentation. While reading work zone signs, composing a detector and text spotter via crop-scaling improves performance (+14.2% 1-NED). Composing work zone detections to provide context further reduces hallucinations (+3.9 SPICE) in VLMs. We predict navigational goals and compute drivable paths from work zone videos. Incorporating road work semantics ensures 53.6% goals have angular error (AE) < 0.5 (+9.9%) and 75.3% pathways have AE < 0.5 (+8.1%).
Traditional animation production decomposes visual elements into discrete layers to enable independent processing for sketching, refining, coloring, and in-betweening. Existing anime generation video methods typically treat animation as a distinct data domain different from real-world videos, lacking fine-grained control at the layer level. To bridge this gap, we introduce LayerAnimate, a novel video diffusion framework with layer-aware architecture that empowers the manipulation of layers through layer-level controls. The development of a layer-aware framework faces a significant data scarcity challenge due to the commercial sensitivity of professional animation assets. To address the limitation, we propose a data curation pipeline featuring Automated Element Segmentation and Motion-based Hierarchical Merging. Through quantitative and qualitative comparisons and user study, we demonstrate that LayerAnimate outperforms current methods in terms of animation quality, control precision, and usability, making it an effective tool for both professional animators and amateur enthusiasts. This framework opens up new possibilities for layer-level animation applications and creative flexibility. Our code is available at https://layeranimate.github.io.
Chirality information (i.e. information that allows distinguishing left from right) is ubiquitous for various data modes in computer vision, including images, videos, point clouds, and meshes. While chirality has been extensively studied in the image domain, its exploration in shape analysis (such as point clouds and meshes) remains underdeveloped. Although many shape vertex descriptors have shown appealing properties (e.g. robustness to rigid-body transformations), they are often not able to disambiguate between left and right symmetric parts. Considering the ubiquity of chirality information in different shape analysis problems and the lack of chirality-aware features within current shape descriptors, developing a chirality feature extractor becomes necessary and urgent. Based on the recent Diff3F framework, we propose an unsupervised chirality feature extraction pipeline to decorate shape vertices with chirality-aware information, extracted from 2D foundation models. We evaluated the extracted chirality features through quantitative and qualitative experiments across diverse datasets. Results from downstream tasks including left-right disentanglement, shape matching, and part segmentation demonstrate their effectiveness and practical utility. Project page: https://wei-kang-wang.github.io/chirality/
NEC is the leading ICT technology provider in the B-to-B market and is actively integrating cutting-edge technologies into its business solutions to drive innovation, enhance capabilities, and create new value for its customers in a broad spectrum of industrial segments. And the recent business focus of NEC is to support digital transformation of business processes of customer enterprises by leveraging technical capabilities in AI, Cyber Security and Communication. This keynote discusses the specific role of NEC's Research in such a business context by sharing a variety of generative and multimodal AI-related use cases that are aimed at solving critical customer challenges in the real-world. From the multimedia perspective, the topics will include world-leading facial recognition technology for security boost and enhanced customer experience, development of drive-recorder video analytics for insurance adjusters leveraging visual language model (VLM) and medical document generation AI service for genuinely supporting overworked clinical doctors. Meanwhile, distributed acoustic sensing technology using optical fiber cables is opening a new opportunity for infrastructure and incident monitoring solutions after integration with AI and ML algorithms. As the common denominator, our commitment of solving critical customer challenges requires (and justifies) nurturing both world-class excellence in performing academic research and accumulated experience and/or culture of application-oriented technology refinement as well as technology combination to ensure business-ready practicality. Also, being the industrial research organization, we are engaged at the forefront of customer co-creation and co-design that play an indispensable role in pinpointing customer's critical challenges. These expertise and practices are indeed the core ingredients of NEC's Research for creating new business opportunities from the technology innovation approach. Furthermore, we also envision that such an industrial lab model in the Generative AI era will become the driver of a new technology paradigm - industry segment-oriented customizable foundation models and business transforming Agentic AI framework.
Intelligent Document Processing (IDP) is critical for unlocking actionable insights from the vast volume of unstructured documents like invoices and medical reports, yet its promise is often unfulfilled as its implementation is typically hindered by significant technical barriers. Traditional IDP systems require deep expertise in programming, machine learning, and intricate model fine-tuning, creating a dependency on specialized data science teams. This effectively sidelines domain experts-the very individuals who possess the critical contextual understanding of the documents-thereby limiting the agility and accuracy of workflow automation. This paper introduces IDPFlow, a novel framework to unify a no-code, user-centric interface with a sophisticated, tool-augmented agentic architecture for end-to-end multimodal document processing, empowering experts such as business analysts and legal professionals to independently build and deploy sophisticated workflows without writing any code. IDPFlow is built upon a powerful agentic architecture, which intelligently utilize a versatile toolkit to execute a range of sophisticated IDP tasks. This toolkit enables a spectrum of high-precision IDP tasks such as multi-class document classification, Document visual question answering (Doc-VQA), key information extraction from text, tables, and checkboxes and long-document summarization. The core of IDPFlow is its dynamic agentic workflow, which redefines user interaction. Upon document upload, the agentic system instantly analyzes the content, classifying sub-documents and proactively suggesting a comprehensive data schema relevant to the use case, shifting the user's role from workflow builder to supervisor. This initial workflow is not static, it can be refined in real-time through simple, conversational instructions, enabling true business agility. Furthermore, the agentic intelligence extends to reusability, allowing existing workflows to be intelligently adapted for new, related tasks, dramatically reducing development time for subsequent use cases. For particularly complex tasks involving long or dense documents, the agentic system can leverage a specialized Multimodal Retrieval-Augmented Generation (MMRAG) pipeline to overcome the context window limitations of standard LLMs. This pipeline utilizes the ColPali model, which excels at generating unified multimodal embeddings, ensuring robust and accurate information retrieval from both textual content and embedded images or diagrams. To foster user trust and ensure verifiability, IDPFlow incorporates a grounded traceback citation mechanism that automatically highlights the precise document segments from which the agent derived its responses, making all outputs transparent and easily auditable. We highlight three key advantages of the framework: 1) Accessibility via an intuitive interface for domain experts; 2) Deep Adaptability and Reusability through dynamic agentic refinement and extensible tools; and 3) Trustworthiness rooted in a verifiable RAG pipeline and granular citation. The framework is projected to reduce end-to-end workflow creation time by 60-70% compared to traditional methods. Its unique combination of a no-code interface and a tool-augmented agentic architecture bridges the gap between technical complexity and domain expertise, accelerating the deployment of powerful, transparent, and scalable IDP solutions across industries.
The rapid expansion of multimedia services, such as video streaming, video conferencing, virtual reality, and cloud gaming, makes maintaining and evaluating high perceptual visual quality essential for user experience and system competitiveness. However, visual content can degrade at multiple stages, including acquisition, compression, transmission, enhancement, and display, where suboptimal enhancement may also introduce artifacts and reduce perceived quality. The core challenge is to reliably measure and predict this perceived quality so that it can be maintained or improved. Perceptual Visual Quality Assessment (PVQA) addresses this by evaluating visual quality from the perspective of human subjects, through subjective studies and objective prediction models. Beyond humans, recent work also extends PVQA to machines and robots, where the goal is to preserve downstream task performance (e.g., segmentation accuracy and planning success) under distortions or bandwidth constraints. This tutorial provides a concise, practice-oriented overview of PVQA: fundamentals and human vision considerations; image and video quality assessment; methods for immersive/3D media; opportunities and challenges in the era of foundation models and GenAI; perceptual optimization loops that close the gap between assessment and decisions in coding, streaming, and embodied perception; and domain applications. Finally, we summarize the key concepts, toolchains, and future opportunities for PVQA to be used in modern multimedia communication.
Major Depressive Disorder (MDD) is a prevalent and severe psychiatric disorder, and its detection remains challenging due to the complexity and variability of its symptoms. Traditional single-modality methods often fail to capture the full spectrum of depressive cues, which has led to the rise of multimodal methods. The ACM Multimedia 2025 ''Multimodal Personality-Aware Depression Detection Challenge'' (MPDD 2025) aims to advance the development of more accurate depression detection models by incorporating multimodal data. In this paper, we proposed a Multi-Level Segment Fusion Based on Adaptive Time-Window Selection (MSF-ATS) method for the MPDD-Elderly Track. To address the challenge of sparse and transient depressive symptoms, we fuse segment-level classifications to obtain subject-level classifications. An adaptive time-window selection based on mean class variance is employed to choose the window with the smallest variance for more stable detection results. Our method achieved an average score of 0.8576 on the MPDD 2025 official test set, significantly outperforming the baseline score of 0.6675.
Document Visual Question Answering (Document VQA) faces significant challenges when processing long documents in low-resource environments due to context limitations and insufficient training data. This paper presents AdaDocVQA, a unified adaptive framework addressing these challenges through three core innovations: a hybrid text retrieval architecture for effective document segmentation, an intelligent data augmentation pipeline that automatically generates high-quality reasoning question-answer pairs with multi-level verification, and adaptive ensemble inference with dynamic configuration generation and early stopping mechanisms. Experiments on Japanese document VQA benchmarks demonstrate substantial improvements with 83.04% accuracy on Yes/No questions, 52.66% on factual questions, and 44.12% on numerical questions in JDocQA, and 59% accuracy on LAVA dataset. Ablation studies confirm meaningful contributions from each component, and our framework establishes new state-of-the-art results for Japanese document VQA while providing a scalable foundation for other low-resource languages and specialized domains. Our code available at: https://github.com/Haoxuanli-Thu/AdaDocVQA.
Multimodal knowledge graphs often separate easily represented information (text) from that which is not (multimedia documents like images, videos, or audio). This severely limits query expressiveness, as the engines lack access to the node contents stored externally. We present MeGraS, the MediaGraph Store, a novel storage and query engine for multimodal knowledge graphs. By storing multimedia documents directly in the graph, MeGraS allows the query engine to leverage their content for enhanced capabilities, making it natively capable of performing operations such as k-NN, segmentation, or deriving non-materialized relations based on visual features. To demonstrate this, we incorporate and extend the pattern-matching query language SPARQL, resulting in a unified framework for storing and managing multimodal knowledge graphs with advanced expressiveness. MeGraS is available as open-source software: http://megras.org
Open-source foundation models are essential for advancing music audio understanding and ensuring access to general-purpose representations for music information retrieval. To this end, we present OMAR-RQ, a model trained with self-supervision via masked token prediction using a large-scale dataset with over 330,000 hours of music audio. We experiment with various input features and quantization options, outperforming existing open self-supervised models in music tagging, pitch estimation, chord recognition, beat tracking, segmentation, and difficulty estimation. Finally, we release our training and evaluation pipelines and model weights at https://github.com/mtg/omar-rq.
Existing video analysis models often lack explainability, perform poorly on long videos, and frequently hallucinate. Commercial solutions are closed-source and costly. We introduce CReLeRI, an open-source system for action detection in untrimmed videos. CReLeRI segments videos using scene and action transitions, detects actions and their arguments and grounds them in 3D space to improve interpretability and reduce hallucinations. The system promotes transparency and trust in AI-driven analysis of complex, real-world videos. A demonstration video is also available.
As numerous edge devices start implementing intelligent components, the challenges of energy consumption, bandwidth efficiency, and privacy gain significance. One proposed solution relies on the paradigm of split inference, which optimizes the delegation of the computational load between edge and remote devices. We developed and implemented the standard-compliant split inference system with an encoder and decoder capable of real-time streaming and processing. Our system outperforms state-of-the-art video compression implementations by an average of 83% bitrate reduction, while preserving privacy. We demonstrate the system's real-time performance on consumer devices, with interactive visualizations of object detection and segmentation, incorporating real-time metrics. Demo video: https://youtu.be/bmCbUo_ZWWU
Reliable, real-time detection of sperm-whale clicks is essential yet difficult in noisy ocean audio streams. We present the first browser-based pipeline that combines self-supervised embeddings with a BiLSTM to label clicks at millisecond resolution. The system attains 99% F1 on the Watkins benchmark, reduces false alarms by 60% against the best published baseline, and analyses one-second segments in ~40 ms on a consumer GPU. An interactive UI overlays multi-algorithm detections on waveform and spectrogram views with drag-zoom and live streaming. Code, pretrained weights and the public demo are released to advance bioacoustic event detection.
In this paper, we present PrivEdit, a zero-shot, interactive image privacy editing system specifically designed for automated sensitive information desensitization. As social networks and smart devices proliferate, the risk of unintended privacy leakage grows, driving demand for personalized, controllable protection tools. PrivEdit is powered by natural-language instructions and integrates a Recognize-Anything model for robust detection and classification of sensitive objects (e.g., faces, license plates, ID cards), followed by GroundingDINO and SAM for high-precision mask extraction. User intents are parsed and disambiguated via GPT-4o, enabling selective target confirmation and iterative refinement. Finally, our editing module performs localized edits-such as adjustable blurring, mosaicking, or replacement via generative editing. With support for multi-round feedback and real-time modification, PrivEdit seamlessly handles both pre-recorded images and live streams, making it ideal for social-media pre-publishing, privacy data desensitization in enterprise or healthcare contexts, and intelligent surveillance applications. By unifying detection, segmentation, intent parsing, and localized editing into one coherent interface, PrivEdit delivers an end-to-end solution for safeguarding visual data. Supplementary materials including the demo video and slides are available at: https://drive.google.com/file/d/13jFBmYgZgxYQLPIAqCeaQhcyzZTHpf7N/view?usp=sharing
This paper presents CrePoster, a data-driven framework to generate aesthetic posters for Chinese cultural relics, aiming to enhance the exhibition experience and promote cultural spread. CrePoster comprises three modules: (1) object segmentation module, (2) content generation module, and (3) poster generation module. Upon processing a cultural relic image, the object segmentation module first leverages a cascaded U2Net-SAM structure to obtain the visual target. Secondly, the content generation module utilizes a multi-target learning-enabled caption generator to produce professional captions. Thirdly, the Multimodal Large Language Model (MLLM) based poster generation module adaptively creates aesthetic parameters, including layout and color scheme, ultimately rendering them into refined posters.
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
PDF ↗The rapid growth of streaming video applications demands multimodal models with enhanced capabilities for temporal dynamics understanding and complex reasoning. However, current Video Question Answering (VideoQA) datasets suffer from two critical limitations: 1) Static annotation mechanisms fail to capture the evolving nature of answers in temporal video streams, and 2) The absence of explicit reasoning process annotations restricts model interpretability and logical deduction capabilities. To address these challenges, we introduce StreamingCoT, the first dataset explicitly designed for temporally evolving reasoning in streaming VideoQA and multimodal Chain-of-Thought (CoT) tasks. Our framework first establishes a dynamic hierarchical annotation architecture that generates per-second dense descriptions and constructs temporally-dependent semantic segments through similarity fusion, paired with question-answer sets constrained by temporal evolution patterns. We further propose an explicit reasoning chain generation paradigm that extracts spatiotemporal objects via keyframe semantic alignment, derives object state transition-based reasoning paths using large language models, and ensures logical coherence through human-verified validation. This dataset establishes a foundation for advancing research in streaming video understanding, complex temporal reasoning, and multimodal inference. Our StreamingCoT and its construction toolkit can be accessed at https://github.com/Fleeting-hyh/StreamingCoT.
Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level and segment-level annotations, comprising over 11,714 segments labeled as Normal or across five Offensive categories: Hateful, Insulting, Sexual, Violence, Self-Harm, along with explicit target victim labels. Our three-stage annotation process yields high inter-annotator agreement (Krippendorff's alpha = 0.817). We propose three tasks to benchmark performance: (1) Trimmed Hateful Video Classification, (2) Temporal Hateful Video Localization, and (3) Online Hateful Video Classification. Results highlight substantial gaps in current models, emphasizing the need for more sophisticated multimodal and temporally aware approaches. The HateClipSeg dataset are publicly available at https://github.com/Social-AI-Studio/HateClipSeg.git.
To meet the growing demand for systematic surgical training, wet-lab environments have become indispensable platforms for hands-on practice in ophthalmology. Yet, traditional wet-lab training depends heavily on manual performance evaluations, which are labor-intensive, time-consuming, and often subject to variability. Recent advances in computer vision offer promising avenues for automated skill assessment, enhancing both the efficiency and objectivity of surgical education. Despite notable progress in ophthalmic surgical datasets, existing resources predominantly focus on real surgeries or isolated tasks, falling short of supporting comprehensive skill evaluation in controlled wet-lab settings. To address these limitations, we introduce WetCat, the first dataset of wet-lab cataract surgery videos specifically curated for automated skill assessment. WetCat comprises high-resolution recordings of surgeries performed by trainees on artificial eyes, featuring comprehensive phase annotations and semantic segmentations of key anatomical structures. These annotations are meticulously designed to facilitate skill assessment during the critical capsulorhexis and phacoemulsification phases, adhering to standardized surgical skill assessment frameworks. By focusing on these essential phases, WetCat enables the development of interpretable, AI-driven evaluation tools aligned with established clinical metrics. This dataset lays a strong foundation for advancing objective, scalable surgical education and sets a new benchmark for automated workflow analysis and skill assessment in ophthalmology training. The dataset and annotations are publicly available in Synapse (https://www.synapse.org/Synapse:syn66401174/files/).