The automated comprehension of complex, multi-modal documents is fundamentally hampered by a disconnect between information extraction and reasoning. Existing systems suffer from inherent limitations. Monolithic models embed reasoning as a black box process, sacrificing transparency and depth. Meanwhile, current agent-based frameworks follow a passive, non-interactive paradigm; they handle static, global inputs rather than information derived from active exploration, which fundamentally restricts their ability to achieve structural understanding and complex reasoning. To bridge this critical gap, we introduce HEAR, a framework for Holistic Extraction and Agentic Reasoning. This innovative framework establishes a synergistic, closed-loop between a deep Vision-Language Model (VLM) driven holistic parsing engine and a collaborative multi-agent reasoning system. Our HEAR initially transforms unstructured documents into a semantically-rich, structured representation, preserving complex layouts and reconstituting multi-page tables. Subsequently, a multi-agent system performs cross-modal analysis, governed by a crucial verification protocol that forces agents to validate findings across textual and visual modalities. A conflict driven re-evaluation mechanism enables the system to dynamically re-engage the document to resolve ambiguities, thereby unifying the perception-cognition cycle. HEAR achieved first place in the ACM MM 2025 Grand Challenge on Large Vision–Language Model Learning and Applications.
论文检索
输入标题、作者或关键词,从 1,620 篇学术成果中精准定位
Video processing and compression are being re-envisioned in the age of AI. Traditional video codecs, which rely on rigid, pre-defined rules, are being augmented and, in some cases, replaced by AI-driven approaches. These new methods leverage machine learning to intelligently analyze video content, allowing for more adaptive and efficient compression. We are going to discuss AOM's new codec AV2, its low- and high-level features that enable significantly smaller file sizes with no perceptible loss in quality, a crucial development for streaming and storage. The shift to AI has also transformed how we evaluate video quality. Traditional metrics, while directionally useful, don't always align with human perception, especially for user generated content (UGC). They fail to capture what's most important for machine vision tasks. We will talk about new AI-based quality metrics that are being developed. They correlate better with a human's subjective experience and a machine's ability to perform tasks like object recognition. Along the way, we'll cover large scale industrial infrastructure challenges and the ways to achieve high reliability and accuracy.
The movement for Sovereign AI is accelerating. Meeting its promise requires vertically integrated AI stacks -spanning data, models, and reasoning systems- that remain sovereign while adhering to shared scientific principles around which global research communities can coalesce. This talk presents BharatGen as a sovereign-yet-shared effort to make AI work for all: creation of datasets, benchmarks, and models that natively support Indian languages, dialects, and code mixing across text, speech, and vision; data pipelines grounded in local realities; and frugal methods that reduce cost and lower barriers. We outline our journey to date across language infrastructure, efficient training and distillation, and early sector pilots. The R&D deep dive will draw from some of our recent work on cross-lingual knowledge distillation for low-resource languages, tokenization/phonetic design for code-mix robustness, or trustworthy document AI with visual grounding focusing on robustness under dialect/code-mix shift, and latency/cost trade-offs. We hope to inspire other Sovereign-AI efforts, especially in the low-resource ecosystems of the Global South and close by inviting international collaborations toward principled research to build people-serving AI.
Data has become the foundation of knowledge, and many companies are growing interested in harnessing AI-based data analysis to unlock its value. The volume of digital data is increasing at an unprecedented pace: market research reports estimate that global data volume, approximately 12.5 zettabytes in 2014, will reach around 180 zettabytes by 2025. Extracting patterns and trends from such big data is crucial for enabling data-driven decision-making. However, a key challenge lies in the enormous computational costs required for large-scale analysis, due to the inherent complexities of the task. Approximate methods are often employed to reduce these costs, but they inevitably trade exactness for efficiency. To overcome this limitation, our research aims to develop a machine learning platform that delivers both speed and accuracy. The core of our platform is computational pruning. This talk will introduce three representative pruning strategies. Specifically, it first introduces a pruning method that uses upper and lower bounds to omit computations. This method efficiently identifies unnecessary processes by using upper and lower bounds of scores to skip unnecessary computations. Next, this talk introduces a method to terminate computations that cannot yield solutions. This method maintains patterns that failed during the search process to avoid repeated futile searches. This talk finally introduces a method that prunes computations through optimistic processing. This method temporarily removes a constraint to find a solution quickly and then verifies if the obtained solution meets the constraint. These strategies can open the path to data analysis techniques that are both efficient and exact, ultimately empowering companies to make more reliable and timely decisions in today's increasingly data-driven world.
In the wave of Artificial Intelligence, along with the proliferation of mobile devices, digital content has been generated and published in an explosive way. Digital content is inherent multimodal: text, image, audio, video, etc. How to effectively and efficiently automate the entire multimodal content lifecycle from idea, to creation and distribution is therefore of great importance. In this talk, we will first delve into the core technological breakthroughs in multimodal content generation, particularly the latest advancements in image and video generation tasks. Then, we will present an agent-based solution from HiDream.ai to interpret user intention and manage content creation-from ideation to final distribution-using three core, interconnected agents. In between, the Content Creation Agent takes a simple user prompt, instantly comprehends the creator's intention, and accurately sources relevant assets from content platform to create multimodal content. Such way frees creators to focus on the story, not the software. The Self-evolving Platform Agent translates global trends and user preferences into strategic directives, guiding models to autonomously generate high-impact content and continuously enriching the platform's ecosystem. The Distribution Agent adapts multimodal content for various social media platforms and analyzes its performance after publishing. Together, these intelligent agents create a seamless ecosystem to connect creators, content and consumers.
The 2024 Nobel Prizes in Chemistry and Physics have once again drawn global attention to AI for Science. The rise of foundation models has further accelerated AI for Science across multiple disciplines. Scientific research is both a touchstone for advancing the intelligence of these models, and the models themselves are accelerators that empower scientific research. As a high-tech enterprise deeply engaged in artificial intelligence, iFLYTEK has in recent years made AI for Science (AI4S) a strategic priority and undertaken a series of initiatives. In this talk, Dr. Xin Li will provide a comprehensive overview of the iFLYTEK Spark Large Language Model and highlight its recent advances. He will outline two principal pathways through which AI accelerates scientific research-deep neural networks and LLM-based approaches-and present iFLYTEK's work along both lines. In addition, he will discuss key challenges and the outlook for the future development of AI for Science. Attendees will gain insights into how AI empowers scientific research and gain inspiration in their own scientific field.
This talk presents research on estimating health status from facial video analysis, focusing on non-invasive approaches to support well-being. The talk introduces robust estimation methods from facial videos for health- and well-being related indicators and real-time systems capable of simultaneously estimating vital signs-including heart rate, respiration, blood oxygen level, blood pressure, and pulse variability. It also explains the feasibility of estimating states such as drowsiness and stress from facial videos. These techniques utilize widely available devices such as smartphones and PCs, enabling continuous and low-burden monitoring in daily life. By demonstrating the feasibility of vision-based health estimation across various contexts, this work highlights its potential for workplace well-being, preventive healthcare, and telemedicine applications. Ultimately, this research aims to develop solutions that leverage cutting-edge technology to promote well-being, contributing to a future where people can lead healthier and more fulfilling lives-both physically and mentally.
The demand for 3D spatial information is rapidly increasing across a wide range of industrial fields. For instance, 3D point cloud data is being actively adopted in sectors such as construction, civil engineering, and disaster prevention to enhance work efficiency and safety. However, 3D point cloud data typically involves extremely large data volumes, which presents a significant challenge; as a result, rapid sharing and utilization over public networks has yet to be fully realized. To address these issues, we are engaged in research and development of compression and transmission technologies for 3D point clouds, as well as contributing to international standardization. In this talk, as part of our initiatives aimed at industrial applications, we will present the latest trends in the international standardization of 3D point cloud compression technologies, such as Geometry-based Point Cloud Compression, and introduce case studies from demonstration experiments utilizing these technologies.
The integration of AI/ML technologies into medical imaging is revolutionizing radiology, offering transformative benefits in clinical workflows. AI-powered Software as a Medical Device (SaMD) solutions not only reduce workload and optimize image interpretation but also unlock critical insights previously undetectable by human eyes-catching the unseen and enabling earlier, more accurate diagnoses. Lung cancer, the leading cause of cancer-related mortality worldwide, is often diagnosed at a late stage, when curative treatment is no longer viable. Early detection is paramount. Traditional screening methods rely heavily on nodule size and growth as indicators of malignancy. However, these criteria alone are insufficient for identifying cancer at its earliest, most treatable stage. eyonis® LCS, the flagship clinical development program of Median Technologies, represents a next-generation AI/ML-based SaMD designed specifically for lung cancer screening [1][2]. It combines Computer-Aided Detection (CADe) and Computer-Aided Diagnosis (CADx) [3][4] capabilities to support clinicians in identifying malignant nodules with greater precision. By leveraging specific architectural choices and deep learning models, eyonis® LCS enhances diagnostic accuracy beyond the current standard of care [5], offering a paradigm shift in early lung cancer detection [6]. This presentation will delve into some of the architectural foundations of eyonis® LCS, highlight its clinical impact, and demonstrate how it empowers radiologists to diagnose lung cancer when patients can still be cured. Through this pioneering technology, Median Technologies is redefining the future of cancer screening and patient outcomes.
The rise of generative AI has democratized media creation, bringing huge promise but also possible perils. While this may seem like a new problem, the generation and manipulation of media has a long history that predates the current AI boom. I'll discuss key insights from our multi-year analysis of content that people shared online. Looking at manipulations such as deepfakes and cheapfakes, as well as misleading contextual manipulations, I'll reveal surprising statistics that challenge common assumptions about the most prevalent types of problematic media. I'll then explore mitigation strategies, including ways to improve information literacy tools, the opportunities and limitations of using AI to detect manipulated content, and how provenance methods paired with AI can help address out-of-context manipulations. Finally, I'll introduce an AI-based tool that can provide additional context for the media we encounter online every day.
NEC is the leading ICT technology provider in the B-to-B market and is actively integrating cutting-edge technologies into its business solutions to drive innovation, enhance capabilities, and create new value for its customers in a broad spectrum of industrial segments. And the recent business focus of NEC is to support digital transformation of business processes of customer enterprises by leveraging technical capabilities in AI, Cyber Security and Communication. This keynote discusses the specific role of NEC's Research in such a business context by sharing a variety of generative and multimodal AI-related use cases that are aimed at solving critical customer challenges in the real-world. From the multimedia perspective, the topics will include world-leading facial recognition technology for security boost and enhanced customer experience, development of drive-recorder video analytics for insurance adjusters leveraging visual language model (VLM) and medical document generation AI service for genuinely supporting overworked clinical doctors. Meanwhile, distributed acoustic sensing technology using optical fiber cables is opening a new opportunity for infrastructure and incident monitoring solutions after integration with AI and ML algorithms. As the common denominator, our commitment of solving critical customer challenges requires (and justifies) nurturing both world-class excellence in performing academic research and accumulated experience and/or culture of application-oriented technology refinement as well as technology combination to ensure business-ready practicality. Also, being the industrial research organization, we are engaged at the forefront of customer co-creation and co-design that play an indispensable role in pinpointing customer's critical challenges. These expertise and practices are indeed the core ingredients of NEC's Research for creating new business opportunities from the technology innovation approach. Furthermore, we also envision that such an industrial lab model in the Generative AI era will become the driver of a new technology paradigm - industry segment-oriented customizable foundation models and business transforming Agentic AI framework.
Surveillance systems such as bodycams and drones often operate under bandwidth constraints that limit video quality and degrade both human monitoring and AI-based analytics. Traditional compression techniques introduce artifacts that obscure critical details, especially in high-motion scenarios. We present a generative AI-powered video compression framework developed by Small Pixels, a spin-off of the University of Florence, designed to deliver Full HD video at significantly reduced bitrates. The system combines edge-side preprocessing for compression resilience with real-time receiver-side super-resolution, enabling up to 50% bandwidth savings while preserving perceptual quality and detection accuracy. Objective evaluations on EgoSeg and VisDrone datasets show +6.2 VMAF improvement and stable YOLOv11 detection performance with 30% less bitrate. Live trials in Singapore, within the Singapore Hatch-X Global Innovation Program, validated real-time operation with minimal latency, demonstrating clearer faces and motion in challenging conditions. The solution integrates seamlessly into existing infrastructures without hardware upgrades, offering a practical path to reliable, high-quality video streaming in bandwidth-limited environments.
The evolving media landscape increasingly demands immersive, non-linear formats supported by innovative tools for content creation and distribution. The Horizon Europe XReco project addresses this need by providing a unified, data-driven ecosystem for next-generation media production, with a focus on extended reality (XR) and virtual production. The platform integrates ingestion of diverse media types (text, images, audio, video, 3D), cross-modal search, 3D content creation, sharing, and monetization. Central to XReco is a metadata-driven ingestion system that overcomes archive fragmentation by enabling efficient organization and access to content from sources like broadcasters, online news, and open repositories. This capability was demonstrated through a short TV documentary on Guglielmo Marconi, created using historical materials assembled via the XReco platform. The platform's Orchestrator module empowers users with powerful cross-modal semantic search capabilities, leveraging neural descriptors to enable queries across different media formats. Editorial teams can retrieve relevant contents searching by keywords like ''telegraph'' or perform reverse image searches to identify and contextualize visual assets like images and 3D models. This unified search functionality significantly enhances content discovery and reuse. A major innovation of the platform consists in providing a set of tools for enhancing the quality of the ingested contents, as well as generating 3D models from 2D assets using state-of-the-art techniques (video super resolution, blind face restoration, NeRF, Gaussian Splatting, Structure from Motion). These services are accessible and tunable via a unified interface, which provides a streamlined user experience and hides the complexity of the underlying technologies. For the Marconi documentary, detailed 3D models of key technological artifacts were created, enabling viewers to interactively explore these objects from multiple perspectives. XReco also supports seamless integration with third-party tools to enrich production workflows. Our documentary incorporated photorealistic digital avatars created with Unreal MetaHuman, animated via motion capture, and featured holoported human experts alongside real presenters within dynamic virtual environments. A noteworthy example is the virtual reconstruction of the RAI Radio Museum in Turin based on Gaussian Splatting, in which avatars from remote locations are developed with Unity and rendered using 4D Gaussian Splatting and Free Viewpoint Video (FVV) technologies. Compatibility with platforms such as Unity and Unreal Engine further facilitates the creation of visually compelling XR experiences. In summary, the XReco platform represents a robust end-to-end solution that effectively tackles the technical and commercial complexities of modern XR and virtual production, paving the way for innovative storytelling in the evolving media ecosystem. During the demo, attendees will have a walkthrough of the platform functionalities, highlighting key technologies for content search, filtering, and processing. They will also be able to enjoy a short documentary about Guglielmo Marconi, produced by our editorial team using XReco technology. After the walkthrough, attendees will have the opportunity to interact directly with the XReco platform to explore its features hands-on-such as testing the search capabilities, creating 3D assets, and experimenting with other available tools. This will provide a more engaging and comprehensive experience of the demo's functionalities. Link to the video: https://drive.google.com/drive/folders/15XTkg-x1U62hQ2dRo2ABiT3LG94CJXO5
Intelligent Document Processing (IDP) is critical for unlocking actionable insights from the vast volume of unstructured documents like invoices and medical reports, yet its promise is often unfulfilled as its implementation is typically hindered by significant technical barriers. Traditional IDP systems require deep expertise in programming, machine learning, and intricate model fine-tuning, creating a dependency on specialized data science teams. This effectively sidelines domain experts-the very individuals who possess the critical contextual understanding of the documents-thereby limiting the agility and accuracy of workflow automation. This paper introduces IDPFlow, a novel framework to unify a no-code, user-centric interface with a sophisticated, tool-augmented agentic architecture for end-to-end multimodal document processing, empowering experts such as business analysts and legal professionals to independently build and deploy sophisticated workflows without writing any code. IDPFlow is built upon a powerful agentic architecture, which intelligently utilize a versatile toolkit to execute a range of sophisticated IDP tasks. This toolkit enables a spectrum of high-precision IDP tasks such as multi-class document classification, Document visual question answering (Doc-VQA), key information extraction from text, tables, and checkboxes and long-document summarization. The core of IDPFlow is its dynamic agentic workflow, which redefines user interaction. Upon document upload, the agentic system instantly analyzes the content, classifying sub-documents and proactively suggesting a comprehensive data schema relevant to the use case, shifting the user's role from workflow builder to supervisor. This initial workflow is not static, it can be refined in real-time through simple, conversational instructions, enabling true business agility. Furthermore, the agentic intelligence extends to reusability, allowing existing workflows to be intelligently adapted for new, related tasks, dramatically reducing development time for subsequent use cases. For particularly complex tasks involving long or dense documents, the agentic system can leverage a specialized Multimodal Retrieval-Augmented Generation (MMRAG) pipeline to overcome the context window limitations of standard LLMs. This pipeline utilizes the ColPali model, which excels at generating unified multimodal embeddings, ensuring robust and accurate information retrieval from both textual content and embedded images or diagrams. To foster user trust and ensure verifiability, IDPFlow incorporates a grounded traceback citation mechanism that automatically highlights the precise document segments from which the agent derived its responses, making all outputs transparent and easily auditable. We highlight three key advantages of the framework: 1) Accessibility via an intuitive interface for domain experts; 2) Deep Adaptability and Reusability through dynamic agentic refinement and extensible tools; and 3) Trustworthiness rooted in a verifiable RAG pipeline and granular citation. The framework is projected to reduce end-to-end workflow creation time by 60-70% compared to traditional methods. Its unique combination of a no-code interface and a tool-augmented agentic architecture bridges the gap between technical complexity and domain expertise, accelerating the deployment of powerful, transparent, and scalable IDP solutions across industries.
We demonstrate an end-to-end system for real-time, multimodal industrial anomaly detection (IAD), built upon a custom hardware platform for synchronized 2D and 3D data acquisition. Our core contribution is a novel cross-modal residual mechanism that identifies defects by quantifying predictive errors between visual and geometric feature spaces. Instead of traditional concatenation, our dual-stream architecture mutually predicts features across modalities, leveraging the prediction residual's magnitude as a direct and robust anomaly indicator. The entire system achieves sub-second inference from acquisition to decision, enabled by efficient depth map analysis that circumvents the complexity of direct point cloud processing, offering a deployable solution for high-speed inspection.
In this industry demonstration, we present SOMIN - an Explainable AI and LLM Platform for Real-Time, Data-Driven Digital Marketing Strategy recommendation. This is achieved through two primary, interconnected subsystems: a predictive model for performance analysis and an LLM for semantic understanding. The system is built on SoWide-ViT, a ''wide and deep'' neural network for advertiser-side CTR prediction. It processes tabular features (campaign settings), text (ad copy, headlines), and visuals (images or keyframes) through TabTransformer, multilingual BERT, and a Vision Transformer (ViT). Replacing the previous ABN model with ViT improved performance, reaching an F1-score of 0.78. Explainability comes from ViT's self-attention, which produces heatmaps highlighting influential regions. High-performing ads showed relevant cues (e.g., gaming objects), while low-performing ones exposed distracting elements. Heatmaps show where the model focuses, while the LLM explains why and suggests improvements. A multimodal GPT-5 interprets creatives, finding flaws such as weak hierarchy, poor CTAs, or off-brand imagery. At scale, the Content Library and Perspective Studies classify competitor ads into 16 marketing concepts, clustering them into Personas (e.g., ''Chocolate Connoisseur'') and Tensions (e.g., ''Work Stress''). The SOINSPIRE module converts Personas and Tensions into insights and campaign propositions, applying marketing theories to generate repeatable, data-driven Expressions.
Recent industry discussions, led by Meta CEO Mark Zuckerberg, envision a radical future where advertisers provide only goals and budgets, while AI systems autonomously manage the entire advertising pipeline - creative, targeting, optimization, and measurement. Zuckerberg described this as a ''redefinition of the category of advertising,'' promising speed, scalability, and simplicity. However, critics warn of risks to transparency and trust, as such systems could displace agencies and in-house teams, leaving brands with little oversight of the ''black box'' outcomes. Alternative approaches frame AI in marketing as a collaborator. Google's ACAI, built with DeepMind and Oxford, treats AI as a co-creator: advertisers provide assets and insights, which are synthesized into structured ''super-prompts'' to guide campaigns, ensuring agency and transparency. Meta's AI Sandbox takes a similar path, offering generative tools for copy, visuals, and resizing while preserving human input. Evidence shows such co-created campaigns deliver higher returns on ad spend, highlighting the value of AI-human collaboration. A third path, seen in platforms like SOMIN, Nielsen, and GWI, emphasizes insight and explainability over automation. Rather than generating ads, these systems analyze successful campaigns to reveal audience, creative, and optimization patterns, embedding AI as a data-driven guide to human creativity. Together, these models illuminate three diverging philosophies: (i) replacement, where AI supplants creative and strategic functions (Meta); (ii) empowerment, where AI scaffolds human creativity (Google); and (iii) insight, where AI explains and enhances strategic thinking (SOMIN and peers). AI adoption is reshaping agencies: while tasks like resizing, reporting, and A/B testing are automated, their real value lies in guiding AI to ensure brand alignment, cultural nuance, and authenticity. This shift demands hybrid skills that merge storytelling with data fluency, AI literacy, and prompt engineering. To adapt, agencies are investing in large-scale upskilling - WPP alone delivered over 150,000 AI training sessions in 2025 - and new roles such as creative technologists, AI trainers, and ethics officers are emerging, positioning agencies as curators of AI-driven creativity rather than passive executors. Hybrid agencies combining storytelling, data science, and engineering - like Accenture Song with its multi-billion AI investments - are disrupting traditional models by offering end-to-end capabilities once spread across multiple firms. Analysts call them ''talent magnets,'' attracting professionals eager to work across disciplines. Meanwhile, Meta's automation push threatens traditional roles, forcing agencies to compete through culture, talent, and authenticity, positioning themselves as AI facilitators that keep human creativity central.
We present MedAI Hub, an integrated multimodal medical platform designed to bridge clinical practice and research by transforming patient-doctor interactions into structured scientific data. This platform supports comprehensive management of multimodal medical records-including clinical notes, medical images, and patient-reported outcomes-while implementing privacy-preserving data sharing mechanisms. Building upon this infrastructure, we introduce two novel AI-driven modules:(1) ITERATE (Image-Text Enhancement, Retrieval, and Alignment): An evolutionary algorithm inspired by Visual Genome that optimizes medical image-text alignment through iterative cross-modal refinement. Leveraging LLM-guided ''DNA evolution'' and multimodal feedback, ITERATE enhances ultrasound image quality for diagnostic tasks, achieving 3.5-7% accuracy gains on ScienceQA and ARC-Easy benchmarks.(2) MedQuery: A graph-driven literature retrieval system that constructs multimodal knowledge graphs from medical literature (text, figures, tables). By aligning PubMed documents with complex clinical queries through semantic relationship modeling, it achieves >90% answer quality win rates and 13-36% accuracy improvements on PubMedQA and MedInquiry datasets. MedAI Hub demonstrates that synergistic integration of clinical data platforms with evolutionary vision-language optimization and multimodal knowledge graphs significantly advances medical AI capabilities, enabling more accurate diagnostics and research insights. The platform and algorithms are publicly available to accelerate innovation in medical AI.
With the wide deployment of visual capture systems, video content restoration plays a key component in the video processing pipeline. We will review the main video content restoration technologies developed in the past decades, including both conventional and deep learning-based solutions. We will highlight the limitations of existing methods, discuss the challenges for the real-world captured video use cases, and point out the potential future research directions.
Balancing energy efficiency and occupant comfort in building HVAC systems is a critical challenge. While generative AI shows promise, its application has been hindered by a reliance on simulations and the inherent instability of its numerical predictions. This paper presents ''Office-in-the-Loop,'' a cyber-physical system leveraging generative AI in a real-world office. Our real-world experiments resolve the energy-comfort trade-off, achieving up to 47.92% energy savings with a 26.36% comfort improvement. We introduce a novel prompting technique, ''Data-Driven Reasoning,'' which compels the AI to justify its predictions with data. This simple addition improves prediction accuracy within ±0.5°C from 50% to 92.31%, paving the way for reliable, AI-driven building automation.