论文检索

输入标题、作者或关键词,从 11,272 篇学术成果中精准定位

会议来源 已选 1 项

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

已选择 1 个会议
支持跨会议组合检索,PDF 均跳转至官方来源
已筛选 AAAI
11,272篇论文
第 4 / 564 页

Francesco Doria, Francesco Percassi, Marco Maratea, Mauro Vallati

We present the Traffic Signal Plans Explorer, a framework for visualising and exploring traffic signal plans generated via PDDL+ planning. Designed to support both traffic experts and non-specialists, the tool offers a web-based interface for high-level network analysis and a SUMO-based adapter for detailed simulation. Users can inspect junction settings and link dynamics, and simulate plan execution step by step. The system bridges planning technology with practical traffic control, enhancing the transparency and usability of automatically generated solutions.

Dinh-Truong Do, Hoang-An Trieu, Van-Thuy Phi, Le-Minh Nguyen, Yuji Matsumoto

Scientific research articles, typically distributed in PDF format, contain valuable knowledge but remain challenging to convert into structured datasets due to fragmented workflows that separate parsing, annotation, and visualization. Existing annotation platforms operate on plain text, which requires an additional PDF-to-text conversion step before annotation, while PDF parsing tools lack automated annotation suggestions. To bridge this gap, we introduce Docora, a system that unifies PDF parsing, automated annotation assistance, and multi-view visualization into a single interactive platform. Docora enables researchers to configure entity and relation schemas for any domain, automatically generates initial annotations using rule-based, model-based, or LLM-based extractors, and provides synchronized visualizations across PDF, text, and graph views. Users can refine annotations directly on the PDF canvas, ensuring consistency between document layout and structured representations. The system’s source code is publicly available to facilitate further research and development.

Alla Chepurova, Aydar Bulatov, Mikhail Burtsev, Yuri Kuratov

Knowledge Graphs (KGs) provide structured, verifiable representations that ground facts and supply large language models (LLMs) with reliable real-world information. Building high-quality KGs from open-domain text remains difficult due to redundancy, inconsistency, and lack of ontology grounding. We present Wikontic, a pipeline that extracts triples from text with LLMs and refines them through ontology-based typing, schema validation, and entity deduplication, yielding compact and coherent graphs. Unlike prior frameworks that lack ontology grounding or perform only partial deduplication, Wikontic uniquely integrates entity canonicalization, alias tracking, and automatic enforcement of Wikidata’s ontology, enabling robust schema-aware construction without manual schema design. Its web interface lets users upload text, visualize graphs, and perform multi-hop question answering. By combining LLM flexibility with Wikidata’s ontological rigor, Wikontic transforms ambiguous text into structured, interpretable, and actionable knowledge.

Yongyang Cheng, Boqin Qin, Zhao Hui, Xu Chen, Tao Zhang, Shang Sun, Haiquan Kang, Xiaojie Xu, Junwei Lv, Lei Yang 等

We present PHOTONS (Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views), a real-time framework for novel view synthesis without requiring camera calibration. Our method reconstructs consistent 3D Gaussian point clouds and synthesizes 2K photo-realistic novel views from arbitrary numbers (>=2) of freely placed cameras. PHOTONS faithfully renders dynamic human bodies amid complex backgrounds, including interactive object manipulation and fine-grained details (e.g., hair strands), while maintaining 25 FPS throughput on commodity GPU like NVIDIA RTX 4090. By combining pose-free spatial point cloud reconstruction with Gaussian parameter estimation, our method demonstrates strong resilience to occlusions and camera perturbations. Additionally, we develop a 3D stereo system that drastically reduces setup complexity compared to existing solutions. Experiments on public and custom datasets show that PHOTONS outperforms state-of-the-art methods in both efficiency and visual quality.

Pi-Wei Chen, Jerry Chun-Wei Lin, Barış Fahri Kahrıman, Zih-Ching Chen, Rafał Cupek, Marek Drewniak

Event detection is essential for surveillance, particularly in retail loss prevention, where accurate and timely monitoring is critical. Vision Language Models (VLMs) provide strong generalization but are inefficient at processing full video streams and are prone to hallucinations induced by redundant frames. We present SmartEyes, a plug-and-play system for real-time retail surveillance. SmartEyes introduces the Perception Cognition Focusing (PCF) framework, which combines lightweight perception with semantic triggering to isolate two keyframes (customer contact and departure) and constrains the VLMs to a focused differencing task. This design reduces hallucination by 44% compared to vanilla VLMs. From the demonstrated retail application, the proposed perception-to-reasoning pipeline is general and directly extends to industrial environments that require reliable event detection and real-time decision-making. Our demo includes a user-friendly Region of Interest (ROI) selection interface and live CCTV monitoring, producing accurate alerts within 1–2 seconds on a single RTX 4080 GPU. This lightweight framework design enables efficient deployment to broader industrial applications.

Saumya Chauhan, Mila Hong

Supporting children with Autism Spectrum Disorder (ASD) requires highly individualized knowledge. However, critical information is often dispersed across documents such as Individualized Education Plans (IEPs), diagnostic assessments, and caregiver notes. Thus, we propose SHARE (Synthesizing Heterogeneous Autism-support Records into Evidence-based Recommendations), a framework that combines diverse autism-related documents into a concise, actionable set of recommendations for caregivers of children with ASD. Feedback is generated using OpenAI’s large language model API, grounded in user-provided evidence with optional web-based extensions for missing details, and citation-linked. After caregivers attempt and then rate recommendations, SHARE uses a Bayesian bandit algorithm with Upper Confidence Bound (UCB) re-ranking to refine future advice. While previous work mostly focuses on drafting static goals, SHARE additionally combines LLM-generated recommendations, caregiver feedback, and interpretable ranking into a pipeline that can adapt over time.

Megha Chakraborty, Darssan L. Eswaramoorthi, Madhur Thareja, Het Riteshkumar Shah, Finlay Palmer, Aryaman Bahl, Michelle A Ihetu, Amit Sheth

AI-driven education platforms have made some progress in personalisation, yet most remain constrained to static adaptation—predefined quizzes, uniform pacing, or generic feedback—limiting their ability to respond to learners’ evolving understanding. This shortfall highlights the need for systems that are both context-aware and adaptive in real time. We introduce PAL (Personal Adaptive Learner), an AI-powered platform that transforms lecture videos into interactive learning experiences. PAL continuously analyzes multimodal lecture content and dynamically engages learners through questions of varying difficulty, adjusting to their responses as the lesson unfolds. At the end of a session, PAL generates a personalized summary that reinforces key concepts while tailoring examples to the learner’s interests. By uniting multimodal content analysis with adaptive decision-making, PAL contributes a novel framework for responsive digital learning. Our work demonstrates how AI can move beyond static personalization toward real-time, individualized support, addressing a core challenge in AI-enabled education.

Louis Carpentier, Wannes Meert, Mathias Verbeke

Time series anomaly detection has received substantial attention over the past two decades, leading to the development of hundreds of algorithms. However, comprehensively understanding this vast landscape remains challenging, particularly for non-experts and novices. In this demonstration paper, we present InTimeAD, an interactive web application that provides access to more than 30 state-of-the-art time series anomaly detection algorithms. InTimeAD is intended to explore the performance of existing as well as custom anomaly detection models in an interactive, hands-on manner. By lowering the entry bar, we support practitioners overwhelmed by the large number of existing techniques, while providing a platform for researchers to rapidly analyze their novel anomaly detection algorithms.

Ettore Caputo, Sergio Greco, Lucio La Cava

We present ARGUS, an end-to-end Argument Mining (AM) tool that exploits Large Language Models (LLMs) to automatically perform all core AM tasks, i.e., Argument Component Segmentation, Classification, Relation Identification, and Relation Classification. Furthermore, ARGUS builds the corresponding argumentation framework (AF) and seamlessly integrates symbolic solvers to compute extensions and perform formal reasoning. ARGUS is designed to ensure broad flexibility and usability, supporting any open-source or commercial LLMs and symbolic solvers, providing a ready-to-use platform for exploring neuro-symbolic approaches to argumentation in both research and practical applications.

Haritha Ananthakrishnan, Harsha Kokel, Kelsey Sikes, Debarun Bhattacharjya, Michael Katz, Shirin Sohrabi, Kavitha Srinivas

We introduce QueryGym, an interactive environment for building, testing, and evaluating LLM-based query planning agents. Existing frameworks often tie agents to specific query language dialects or obscure their reasoning; QueryGym instead requires agents to construct explicit sequences of relational algebra operations, ensuring engine-agnostic evaluation and transparent step-by-step planning. The environment is implemented as a Gymnasium interface that supplies observations---including schema details, intermediate results, and execution feedback---and receives actions that represent database exploration (e.g., previewing tables, sampling column values, retrieving unique values) as well as relational algebra operations (e.g., filter, project, join).We detail the motivation and the design of the environment. In the demo, we showcase the utility of the environment by contrasting it with contemporary LLMs that query databases. QueryGym serves as a practical testbed for research in error remediation, transparency, and reinforcement learning for query generation.

Suman Adhya, Debarshi Kumar Sanyal

To address the challenge of interpreting evolving themes in temporal text, we present DTECT (Dynamic Topic Explorer & Context Tracker), an interactive, end-to-end system for uncovering thematic dynamics. The system integrates a complete pipeline that supports data preprocessing, multiple model architectures, and dedicated metrics to analyze temporal topic quality. To enhance interpretability, DTECT features LLM-driven automatic topic labeling, trend analysis, interactive visualizations with document summarization, and a natural language chat interface. This cohesive platform empowers users to intuitively explore how topics change over time.

Teresa Zhang

Scaling long-context and agentic LLMs is increasingly limited by memory capacity and bandwidth rather than FLOPs. I propose an algorithmic framework for context engineering that models placement, compression, and scheduling as coupled optimization problems with explicit accuracy-efficiency trade-offs. Concretely, I aim to develop (1) salience-aware retention/eviction policies with provable approximation guarantees relative to an ideal oracle; (2) tier-dependent compression schemes that bound error propagation across memory levels; and (3) probabilistic prefetch/scheduling that controls tail latency. I will evaluate on long-context language modeling and reasoning benchmarks, isolating each component via ablations and comparing against heuristic baselines under controlled bandwidth/capacity regimes. Results target improved throughput and energy metrics at near-baseline quality, advancing principled, hardware-aware inference without requiring custom hardware.

Tianruo Rose Xu

Reasoning-based large language models now often produce natural-language thinking traces alongside their answers, but it remains unclear whether these verbalized uncertainties faithfully reflect their knowledge or can be used to improve factuality. We study this question for long-form, knowledge-intensive biography generation. Our pipeline decomposes thinking traces and responses into atomic facts, filters out planning-style content, labels factual reasoning by certainty, and aligns response facts to their supporting reasoning, enabling plan-based filtering, self-verification, and a classifier that predicts factuality from facts and associated reasoning. Preliminary results suggest that high-certainty reasoning is more likely to be included and correct and that structured use of these signals can improve factual precision, though broader validation across models and dataset will be needed.

Hua Xu

While deep learning excels at decoding neural signals, the opacity of state-of-the-art models limits their scientific utility and clinical trustworthiness. We propose a research that bridges this gap by integrating high-performance architectures—specifically Transformers and Graph Neural Networks—with mechanistic interpretability and neuro-symbolic reasoning. This proposal aims to uncover verifiable mappings between artificial computational circuits and biological dynamics without compromising decoding accuracy. Validated through rigorous benchmarking and wet-lab experiments, this work establishes a foundation for transparent brain-computer interfaces and accelerates fundamental neuroscience research.

Tan Xeng Ian

Low-Rank Adaptation (LoRA) has emerged as a practical and efficient method for fine-tuning large language models under limited computational budgets. However, recent studies have shown that LoRA can suffer from training instability when applied to models with large embedding dimensions, due to the imbalanced in magnitudes between its low-rank matrices. In this work, we propose a novel regularization strategy that stabilizes LoRA training by penalizing logarithmic magnitude differences between the low-rank matrices, showing theoretically that it should lead to efficient feature learning. We further propose evaluation methods to systematically assess training stability and performance of our proposed solution along with other LoRA variants.

Zhifu Wei

Sequential recommendation (SR) aims to model users' dynamic preferences from their historical interaction sequences to provide personalized recommendations. However, data sparsity remains a core bottleneck limiting the performance of sequential recommendation models. Existing mixup methods face two major challenges: 1) They cannot effectively address the data sparsity dilemma in long-tail scenarios. 2) It is difficult to maintain the Semantic structure of augmented samples during the random mixing process. To address these challenges, this study proposes the Semantic-Aware Data Augmentation (SADA) framework, which utilizes large language models (LLMs) to generate semantic embeddings. This framework allows for the fusion of both collaborative and semantic signals, alleviating the representation deficiency of long-tail items. Additionally, through semantic-guided mixup, the framework preserves semantic structure consistency at both the user and item levels, thereby avoiding semantic structure degradation caused by traditional random mixing. This approach is expected to significantly improve recommendation performance and generalization ability across multiple datasets and application scenarios. In a broader context, this research aims to drive the evolution of data augmentation in sequential recommendation from heuristic methods to a semantic-driven paradigm, helping to build more personalized, accurate, and socially valuable recommendation services.

Jarin Tasneem

Federated learning (FL) has rapidly emerged as a pivotal framework for cross-silo collaborative training while keeping sensitive data localized, driven by growing data volumes and major privacy concerns. Within this paradigm, vertical federated learning (VFL) enables collaboration among parties holding different features of the same sample space, powering tasks like fraud detection, medical diagnosis, and credit scoring. However, the participation of multiple entities creates new vulnerabilities to malicious interference. One critical yet underexplored threat in VFL is the Byzantine poisoning attack, where an adversary intentionally corrupts training to degrade overall model performance. This work reveals a practical vulnerability showing how a single malicious participant can significantly reduce inference accuracy in a VFL system by breaking cross-view association through feature-space corruption. Our findings emphasize the urgent need for robust, VFL-specific defenses to ensure reliability in collaborative, cross-silo AI systems.

Jiaen Sun

This research statement proposes to measure and mitigate speaker entanglement, where accent features inadvertently encode who is speaking in accented automatic speech recognition (ASR). We argue that entanglement inflates scores under lenient split for the same speaker and worsens fairness gaps across accents, and we outline a parameter-efficient mitigation that combines adversarial de-speakerization with safe conditioning. The plan is grounded in established results in accented ASR, domain-adversarial learning, and parameter-efficient fine-tuning; it is feasible with public datasets and a frozen Whisper backbone, and can potentially guide low-resource data collection.

Abhiram Srivatsa Kadaba

Modern generative models often violate basic physical principles. Shadows drift, geometry becomes inconsistent across views, and measurement models are ignored, which limits trust in both video synthesis and computational imaging. We propose a finite time Schrödinger Bridge (SB) world model that formulates generation as entropy regularized optimal transport from a simple prior to a distribution that is consistent with both data and physics. Instead of applying consistency corrections only at the final output, the framework introduces geometric and physical structure directly along the generative path. For video, the model enforces multiview geometric constraints through reprojection and epipolar agreement, homographies, and depth guided warping. For imaging, it incorporates differentiable optical operators, including point spread function based defocus models and lightweight Fourier propagation for coherent and partially coherent settings. When camera poses are known, the model penalizes reprojection error and warp aligned photometric or feature inconsistencies. When poses are unknown, a compact motion or flow estimator encourages cycle consistent trajectories. A lightweight UNet or Vision Transformer backbone, together with a short SB horizon, maintains computational efficiency. Evaluation will measure three dimensional and temporal consistency, physics fidelity through forward simulation residuals, and overall generative quality and efficiency using FID, KID, and FVD. Comparisons will include modern video diffusion models, plug and play data consistency methods, and unconstrained SB variants. The central hypothesis is that constraining the entire generative trajectory, rather than only the final frame, can shorten sampling while improving cross view coherence and physical plausibility across diverse sensing modalities, including cameras, microscopes, and medical imaging systems.

Aaron Soh

Recent advances in Large Language Models (LLMs) have achieved state-of-the-art performance in Automatic Speech Recognition (ASR), surpassing ASR-only systems such as Whisper. However, their application to other speech processing tasks, particularly speaker diarisation (SD), remains underexplored. This work proposes extending existing speech-aware LLM architectures with diarisation-specific training and context-based prompting to enable joint transcription and segmentation of multi-speaker audio. By exploiting the semantic reasoning and multilingual capabilities of pretrained LLMs, the proposed approach aims to improve diarisation accuracy, enhancing accessibility for assistive technologies and real-time captioning applications that rely on accurate speaker-aware transcriptions.