论文检索

输入标题、作者或关键词,从 11,272 篇学术成果中精准定位

会议来源 已选 1 项

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

已选择 1 个会议
支持跨会议组合检索,PDF 均跳转至官方来源
已筛选 AAAI
11,272篇论文
第 5 / 564 页

Jinseo Shim

This study is grounded in prior work on program induction framework with a structured latent program space, called Program Lattice Auto Encoder(PLAE). It preserves compositional structure by training an encoder where programs and their compositions correspond to integer linear combinations of program bases, forming a discrete program lattice that captures the geometric structure of compositional reasoning. Based on it, this paper proposes a novel extension of the PLAE aimed at improving generalization and efficiency by choosing a cylindrical lattice latent space instead of plane, which can represent invariant programs. The core hypothesis is that only isometric transformations conserve compositional properties of lattice structure and therefore developable surfaces such as a cylinder or cone are permissible as embedding space. Moreover, through demonstrating a contradiction of lattice on conical manifolds, it conclude that only cylinder is a possible embedding manifold for lattice structure.

Srivarshinee S

This proposal aims to investigate epistemic uncertainty - uncertainty about knowledge or truth, often conveyed by modals like might or probably in Large Language Models (LLMs). By probing how such cues affect reasoning, we seek to achieve controllable epistemic sensitivity: enabling mod- els to interpret and adapt to uncertainty. Using activation- level analyses and multilingual benchmarks, this work ad- vances transparent, context-aware, and trustworthy reasoning in uncertainty-critical domains.

Kin Meng Ng

Large language models (LLMs) have rapidly advanced, but their growing compute demands limit accessibility in under-resourced regions like Southeast Asia (SEA). While hybrid architectures combining Attention and State-Space Models (SSMs) offer efficiency gains, most rely on sequential interleaving, leaving the potential of parallel-head mixing largely under-explored. However, the recent Falcon-H1 family of models has demonstrated that parallel-head hybrid architectures are not only viable, but scalable to state-of-the-art levels. I propose investigating this parallel-head architecture as a foundation for efficient, multilingual SEA LLMs. My short-term goal is to adapt Falcon-H1-1.5B via vocabulary expansion and continuous pretraining, mitigating token fragmentation and enabling low-resource adaptation to 9 SEA languages. In the longer term, I will develop a dynamic token routing mechanism to optimize token-level compute allocation within hybrid layers, aiming to maximize efficiency without sacrificing the expressive power needed for complex multilingual contexts. Evaluation will utilize the SEA-HELM framework to assess whether these parallel-hybrid innovations can democratize access to high-performance AI for SEA communities.

Yimeng Liu

RNA 3D structure prediction is essential for understanding regulatory mechanisms, catalysis, and therapeutic RNA design, yet progress has lagged behind proteins due to limited structural data and the complexity of RNA folding. This work proposes a data-efficient, physics-informed deep learning framework for full atomistic prediction of transfer RNA (tRNA) tertiary structures directly from sequence. Our approach will integrate pretrained RNA embeddings, predicted secondary structure constraints, and SE(3)-equivariant graph attention to model long-range geometric relationships. A two-stage design will first predict global phosphate backbone coordinates, then reconstruct nucleobase atoms using a local geometry-aware decoder. A multi-objective loss will combine geometric accuracy with chemical and biophysical plausibility to enforce valid torsion angles, base-pairing, and steric constraints. We will benchmark against physics-based (VFold) and neural network–based (DeepFoldRNA) models to assess generalization under data scarcity. Ultimately, this framework aims to advance RNA 3D modeling with improved stability, interpretability, and capacity to generalize beyond well-characterized RNA families, supporting future applications in rational RNA engineering and structure-guided RNA design.

Napassorn Litchiowong

I present a compact, testable architecture that endows learning agents with continuous proto-emotional dynamics and interpretable modulators (Persona, Ego, Shadow, Self). The design grounds these modulators in a computational interpretation of Jung’s Map of the Soul, mapping each archetype to a differentiable control that modulates policy selection via a bounded, low-dimensional affect vector. I describe concrete modular implementations, a staged experimental program (toy domains → multi-agent/social tasks → nonstationary transfer), baselines, ablations, and reproducible evaluation metrics.

Xinyu Li

The widespread adoption of artificial intelligence (AI) in cybersecurity has led to the emerging threat of AI-driven cyberattacks, such as LLM-empowered Advanced Persistent Threats (APTs), challenging the effect of conventional deception defense mechanisms. To fill this critical gap, my work aims to develop a game-theoretic defense AI agent capable of providing the optimal deception resource deployment strategy, to establish AI-driven defenses against AI-empowered cyberattacks. In this proposal, I model the attacker and defender interaction as a dynamic game with incomplete information between AI agents, and then derive the equilibrium defense strategies. Synthetic data based experiments and real-world implementations would be conducted to validate the proposed framework. This study has the potential to improve the effectiveness of deception defense in three dimensions: scalability, real-time capability, and strategic intelligence.

Hai Le

Post-training quantization is widely used to compress large language models (LLMs) for efficient deployment in resource-constrained environments. However, recent work shows that quantization, especially aggressive schemes such as 4-bit QLoRA, can substantially degrade safety alignment, making models more vulnerable to harmful completions and jailbreaks. In this work, we investigate these safety risks and propose a mitigation strategy: projecting quantized parameters back into safety-aligned subspaces. First, we empirically measure safety degradation on benchmark datasets using both safety and utility metrics. Next, we explore projection-based restoration methods to recover alignment-preserving directions in the LoRA adapters of quantized models. Finally, we study how quantization affects mechanistic safety neurons and how hybrid-precision designs can preserve them. By foregrounding the safety implications of model compression, this work aims to support more robust, deployment-ready, and ethically aligned LLMs.

Samruddhi Kamble

Sexual trauma leaves wounds that science cannot see, yet survivors live with them every day. Traditional tools rely on words or self-reports, often forcing survivors to "bleed in silence" when their pain is doubted or dismissed. Trauma, however, is not one-dimensional. It disrupts multiple brain networks and produces states fear, vigilance, detachment that cannot be captured by words alone. This creates the need for approaches that reveal trauma's complexity in ways that are both objective and interpretable. Our framework that combines fMRI, EEG, and DINO to investigate machine learning models can identify neural or psychological patterns associated with trauma responses. Instead of producing abstract scans or opaque predictions, the system will generate exploratory measures of trauma response that support therapists' understanding while guiding future research. These measures will be presented through a simple dashboard that summarizes three indices (TPI, DI, RBS) alongside heatmaps and plain language notes. By turning complex data into clear, anonymized session snapshots, the dashboard provides researchers with an output that can be compared across participants and refined in future work.

James Blossom Eleojo

This paper presents an AI-driven framework for real-time reverberation control in dynamic environments. The system integrates parametric modeling in Grasshopper, Pachyderm acoustic simulation, and machine learning to create a closed-loop controller. A CNN estimates reverberation time from audio signals, while a reinforcement learning agent dynamically adjusts panel absorption coefficients to maintain optimal acoustics. Evaluation showed the system should be able to maintain T60 within 0.15 s of the target under varying occupancy and source positions, outperforming static treatments and enabling self-regulating acoustic environments for improved auditory experiences.

Zhiqing Cui

Spatiotemporal forecasting has seen remarkable progress with the advent of deep learning, particularly with Spatiotemporal Graph Neural Networks (STGNNs). These models excel at answering the what question: predicting future numerical values with high accuracy. However, they fail to answer the crucial why question. In high-stakes domains such as meteorology, urban planning, and public health, this opacity creates a critical bottleneck for adoption. A model that predicts a severe pollution event without explaining its atmospheric drivers is a black box, limiting its trustworthiness and utility for decision-makers who need actionable, causal insights. To address this critical gap, I propose a long-term research project to develop Causal-LLM, a new class of foundation models for spatiotemporal data that are both predictively powerful and causally interpretable. My central thesis is that genuine interpretability cannot be an afterthought; it must be designed into the model's core learning process. By adapting the powerful Time-LLM reprogramming framework and introducing a novel training methodology I term causal data synthesis, Causal-LLM will learn to not only forecast future states but also to articulate the human-understandable causal narratives behind them. This research will make two primary contributions: (1) a novel hybrid architecture that synergizes the perceptual power of GNNs with the reasoning capabilities of LLMs for complex physical systems, and (2) a new training paradigm that explicitly teaches this mapping. A successful project would provide a blueprint for a new class of trustworthy foundation models for science, enabling applications such as a climate model that not only predicts a flood but also explains the atmospheric river causing it, empowering authorities to make more informed and trusted decisions.

Yi Khuen Chai

The scaling parameter β in Direct Preference Optimization governs a fundamental trade-off: low β produces weak gradients that fail to learn from ambiguous preferences, while high β amplifies updates and causes excessive drift from the reference policy. Prior work treats β as fixed or scheduled throughout training. We introduce DualLoop-DPO, which modulates β via dual feedback: a fast loop raises β temporarily on high-uncertainty batches to enforce stronger preference margins, while a slow loop uses EMA-smoothed KL tracking to regulate policy drift. Experiments on preference alignment benchmarks show consistent improvements over existing static-β, β-scheduling, and dynamic-β baselines. These findings suggest that dual-loop β control—responding to uncertainty for learning and divergence for stability—offers a promising direction for preference-based fine-tuning.

Krish Bhatnagar

Manic episodes in bipolar disorder are characterized by acute behavioral escalation requiring early intervention. This research proposes a multimodal digital phenotyping framework integrating keystroke dynamics with circadian rhythm features to forecast manic episodes 3-7 days prior to clinical onset. The system leverages a hybrid architecture of temporal convolutional and recurrent neural networks with personalized adaptation. It generates risk predictions and clinically actionable alerts while ensuring user privacy through strict on-device processing and data encapsulation. This framework addresses a critical gap in mental health-care: providing passive, unobtrusive monitoring to detect pre-onset behavioral signatures within a clinically actionable window.

Woo Jon Hou Ainsley

Generative AI shows strong capabilities in language, reasoning, and code but remains prone to hallucinations—outputs that are fluent yet incorrect. In cybersecurity, such errors pose serious risks, from misleading analysts to potential adversarial exploitation. This project investigates hallucinations in three directions: (1) creating benchmarks and interpretability tools to characterize them in security contexts; (2) developing mitigation strategies such as retrieval-augmented generation, symbolic-neural hybrids, and uncertainty-aware decoding; and (3) integrating these methods into real-world workflows like vulnerability assessment, malware analysis, and penetration testing, while exploring how attackers might exploit hallucinations. Evaluation will combine accuracy metrics, human-in-the-loop studies, and red-team simulations. By bridging theory and applied system design, the work aims to advance understanding of hallucinations and improve the reliability of AI in cybersecurity, with broader implications for other high-stakes areas such as healthcare and law.

Tongle Zhao, Fang Fan, Huiqi Zhao, Xiaodu Liu

Encrypted traffic classification has become increasingly important in network security. To address the difficulty of existing architectures in collaboratively modeling spatio-temporal features, we propose BiST-Mamba, a novel dual-branch spatio-temporal Mamba network that enables simultaneous representation of spatio-temporal features. To the best of our knowledge, this is the first work to introduce VMamba into encrypted traffic classification. Preliminary experiments on a small-scale dataset show that our accuracy and F1 scores reach 94.13% and 93.41%, respectively. The method achieves promising classification performance, demonstrating the potential of the model for effective spatio-temporal modeling.

Yuhao Zhang, Ningkang Peng, Yafei Liu, Lin Li, Masaru Kitsuregawa, Yanhui Gu

The two-dimensional (2D) graph structure of a molecule encodes abundant latent property information. A well-designed molecular graph encoder can capture informative low-dimensional dense representations of molecules, which can subsequently be applied to a widerange of downstream tasks. To achieve fine-grained anddiscriminative molecular representations that capture localized structural information, we propose an novel atom-level adaptive receptive field encoder, enabling each atomic node in the molecular graph to dynamically adjust its receptive field size. To the best of our knowledge, we are the first to introduce an effective rank-guided pruning strategy for 2D molecular graphs.

Tianyu Zhang, I-Chao Shen, Haoran Xie

Anime hair design is crucial but challenging, as it conveys personality and emotion through stylized geometry and layered structure. In this work, we propose a sketch-guided approach for intuitive control of multimodal diffusion transformers (MMDiT) to generate semantically consistent anime hairstyles. We adopt a wisp-level flowline input integrated with a fine-tuned MMDiT to transfer hairstyles while preserving character identity. We believe that this fine-grained sketch control within the MMDiT framework may offer a promising path for structured anime hair editing.

Ryan Zhang, Herbert Woisetschläger

Real-world AI systems are tackling increasingly complex problems, often through interactions among Large Language Model (LLM) agents. When these agents develop inconsistent conventions, coordination can break down. Applications such as collaborative coding and distributed planning therefore require reliable, consistent communication, and scalability is a central concern as systems grow. We introduce Schema-Induced Games for Naming (SIGN), a naming game that examines how lightweight structure can steer convention formation. We compare schema-induced communication to unconstrained natural language and find faster convergence with up to 5.8× higher agreement. These results suggest that minimal structure can act as a simple control knob for efficient multi-agent coordination, pointing toward broader applications beyond the naming game.

Rongyu Zhang, Dingyuan Zhang, Haopeng Li

Proactive dialogue systems, which are designed to guide conversations toward predetermined goals. However, contemporary LLMs predominantly function as passive assistants, mechanically executing human instructions. A key challenge contributing to this limitation is the inherent difficulty in acquiring and annotating high-quality training data for proactive dialogue. Consequently, the scarcity of such data results in a notable deficiency in the proactive conversational capabilities of current LLMs.In this paper, we introduce PANDA (Proactive Agent-based Negotiation Dialogue Augmentation), a method designed to generate accurate, complex, and diverse proactive dialogue data for a challenging task—financial dispute mediation—where a LLM acts as the mediator. PANDA leverages a novel self-evolving synthesis process to manage a pool of user profiles and generate dialogues through structured interactions between multiple LLM-driven agents. To ensure data fidelity, we propose a comprehensive evaluation framework and build a two-level validation system combining automated and expert human verification. Our experiments demonstrate that an 8B-parameter model, trained on our synthesized dataset, achieves state-of-the-art results in the task's evaluation framework. Its performance rivals top closed-source models guided by heavily engineered prompts, even when provided with only essential information.

Hanzhang Yuan, Sheng Li

Estimating causal effects under network interference is challenging especially when edges are heterogeneous and nodes share latent dependencies. We study this realistic setting and propose MVDR, a targeted maximum likelihood (TMLE) framework that learns multi-view representations of covariates and exposure on heterogeneous networks while achieving double robustness: consistency holds if either the outcome model or the exposure density is correctly specified. MVDR supports multiple network interventions using only the observed network structure. On three semi-synthetic datasets, MVDR reduces intervention-level prediction error against baselines, and remains stable under misspecification.

Zhenyu Yu, Mohd Yamani Idna Idris, Pei Wang, Rizwan Qureshi

Quantitative remote sensing estimation is critical for environmental monitoring, providing continuous measures of vegetation indices, canopy height, and carbon stock. Traditional radiative-transfer models and empirical regressions require expert knowledge and generalize poorly, while deep learning methods remain task-specific. We propose SatelliteCalculator+, a DINOv3-powered multi-task foundation model for continuous regression of spectral and structural variables. The framework combines prompt-driven cross-attentive adapters with lightweight MLP decoders, enabling efficient dense prediction from frozen features. To overcome limited supervision, we synthesize over one million paired samples from SPOT 6/7 imagery using physically defined formulas. On the Open-Canopy dataset, SatelliteCalculator+ achieves competitive accuracy across eight ecological variables while reducing inference cost, demonstrating the promise of self-supervised transformers and scalable multi-task learning for large-scale Earth observation.