论文检索

输入标题、作者或关键词,从 6,795 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
6,795篇论文匹配“Theory”
第 7 / 340 页

Dong Wang, Jie Jiang, Weidong Min, Lixin Zhan, Xinpeng Zhao, Ze Zhang

Open-vocabulary 3D understanding aims to align 3D representations with a unified vision-language semantic space. However, existing methods suffer from the challenge of structural asymmetry caused by sparse observations and holistic geometries. Additionally, the inherent semantic chasm between discrete coordinates and abstract natural language hinders cross-modal mapping. This work introduces SpecBridge, a 3D-2D-Text pre-training framework that leverages CLIP priors as a foundational bridge to connect three modalities by synergizing spectral graph theory with transitive semantic learning. The method consists of two main sub-modules. First, Spectral Eigen-Modal Alignment is presented to construct cross-modal correlations within the intrinsic eigen-spectral space via Laplacian eigendecomposition. By aligning low-frequency geometric harmonics with CLIP-guided observation inference, it maintains structural consistency between 3D geometries and 2D projections. Second, a Transitive Spectral-Semantic Alignment method is developed to establish a 3D→2D→Text propagation chain across the CLIP priors bridge. It distills dense CLIP priors into 3D representations through transitive distillation, effectively mitigating the semantic chasm between geometry and language. Extensive experiments confirm that the proposed SpecBridge demonstrates state-of-the-art performance by overcoming challenges triggered by modality discrepancies.

Minxi Yan, Yihua Shao, Yanling Pan, Siyu Chen, Hongjuan Pei, Hao Tang, Fei Ma, Jingcai Guo, Nicu Sebe

Test-time scaling (TTS) has demonstrated remarkable potential in enhancing the reasoning capabilities of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs). However, its application has primarily been limited to domains such as mathematics and programming, owing to their reasoning-intensive nature and the ease of result verification. Its utility in other knowledge-intensive fields, such as medicine and general scientific research, remains underexplored. To bridge this gap and unlock the potential of TTS in broader domains, we propose Cross-Domain TTS, a novel framework that enables task-tailored scaling. This framework consists of two key components: a conformal prediction-based cold-start strategy and an information-gain-based dynamic reasoning adjustment. The CP-based cold-start strategy guides the model's initialization during test-time scaling based on conformal prediction theory, while the information-gain-based dynamic reasoning adjustment guides the model's reasoning progress through a progress vector according to the information gain of reasoning steps. We conducted experiments using LLMs and LVLMs on cross-domain benchmarks. Our results demonstrate that the proposed framework consistently improves performance across various domain-specific datasets. For instance, in the medical domain, it achieves an improvement of up to 17% in pass@1 accuracy while reducing inference latency and saving up to 30% in token consumption. Code is available at https://github.com/Yan0613/Cross-Domain-TTS.

Fengming Zhu, Fangzhen Lin

Stochastic games have become a prevalent framework for studying long-term multi-agent interactions, especially in the context of multi-agent reinforcement learning. In this work, we comprehensively investigate the concept of constant-memory strategies in stochastic games. We first establish some results on best responses and Nash equilibria for behavioral constant-memory strategies, followed by a discussion on the computational hardness of best responding to mixed constant-memory strategies. Those theoretic insights are later verified on several sequential decision-making testbeds, including the Iterated Prisoner's Dilemma, the Iterated Traveler's Dilemma, and the Pursuit domain. This work aims to enhance the understanding of theoretical issues in single-agent planning under multi-agent systems, and uncover the connection between decision models in single-agent and multi-agent contexts. The codebase and the full version of this paper is available at github.com/Fernadoo/Const-Mem.

Yan Liu, Zeyu Ren, Pingzhong Tang, Zihe Wang, Yulong Zeng, Jie Zhang

Deterministic auctions are attractive in practice due to their transparency, simplicity, and ease of implementation, motivating a sharper understanding of when they can attain the same outcomes as randomized mechanisms. We study deterministic implementation in single-item auctions under two notions of outcomes: (revenue, welfare) pairs and interim allocations. For (revenue, welfare) pairs, we show a separation in discrete settings: there exists a pair implementable by a deterministic Bayesian incentive-compatible (BIC) auction but not by any deterministic dominant-strategy incentive-compatible (DSIC) auction. For continuous atomless priors, we identify conditions under which deterministic DSIC auctions are equivalent to randomized BIC auctions in terms of achievable outcomes. For interim allocations, under a strict monotonicity condition, we establish a deterministic analogue of Border's theorem for two bidders, providing a necessary and sufficient condition for deterministic DSIC implementability. Using this characterization, we exhibit an interim allocation implementable by a randomized BIC auction but not by any deterministic DSIC auction.

Minhui Zhang, Xuehan Zhao, Xin Zhang, Jiaqi Liu, Zhiwen Yu, Bin Guo

Learning to Defer operates by deferring AI-uncertain samples to humans, enabling the system to outperform either alone. Some works extend it by introducing human-AI combination to deferral. However, they remain limited to binary choices. Given that AI, humans, and human-AI combination are each indispensable, we extend the binary deferral paradigm to a ternary one. There are two challenges: i) how to effectively select among the three modes, especially distinguishing humans from combination? ii) how to design an enhanced combination without being degraded by unreliable model predictions? To address the challenges, we propose TriHAI, a tri-mode deferral method that routes samples to suitable modes based on normalized confidence scores obtained through a discriminative gating network, and enhances combination mode by fusing model predictions refined by human-guided Conformal Prediction and human predictions via Bayesian theory. Experiments on three datasets show that TriHAI surpasses other human-AI baselines by up to 9.57% in accuracy.

Nicholas Teh

We study proportional representation in the temporal voting model, where collective decisions are made repeatedly over time over a fixed horizon. Prior work has extensively investigated how proportional representation axioms from multiwinner voting (e.g., justified representation (JR) and its variants) can be adapted, satisfied, and verified in this setting. However, much less is understood about their interaction with social welfare. In this work, we quantify the efficiency cost of enforcing proportionality. We formalize the welfare-proportionality tension via the worst-case ratio between the maximum achievable utilitarian welfare and the maximum welfare attainable subject to a proportionality axiom. We show that imposing proportional representation in the temporal setting can incur a growing, yet sublinear, welfare loss as the number of voters or rounds increases. We further identify a clean separation among axioms: for JR, the welfare loss diminishes as the time horizon grows and vanishes asymptotically, whereas for stronger axioms this conflict persists even with many rounds. Moreover, we prove that welfare maximization under each axiom is NP-complete and APX-hard, even under static preferences and bounded-degree approvals, and provide fixed-parameter algorithms under several natural structural parameters.

Zhejing Hu, Yan Liu, Zhi Zhang, Sean Fontaine, Gong Chen

The knowing--doing gap, the mismatch between ideas articulated during model reasoning and the realized creative artifact, remains a fundamental challenge in creative AI and persists in LLM-based artistic creation. This paper responds to this gap by introducing a systematic and interpretable framework for examining how user intent is articulated during model reasoning and selectively realized, or lost, during action in LLMs, using symbolic music composition as an analytical lens. We present a realizational process theory that formalizes creative generation and localizes the knowing--doing gap at the realization stage. We instantiate this theory with Music Atelier (Mutelier), an LLM-as-a-Judge framework that operationalizes idea-level realization analysis and makes the gap observable and analyzable in practice. Across diverse evaluation settings, we show that Mutelier reveals reliable, previously invisible failure modes in which intent-aligned ideas are articulated during reasoning but fail to materialize in the final artifact. By reframing artistic creation as a realizational process, this work provides a principled process-level foundation for understanding how knowing--doing gaps emerge in machine creativity.

Woo Jon Hou Ainsley

Generative AI shows strong capabilities in language, reasoning, and code but remains prone to hallucinations—outputs that are fluent yet incorrect. In cybersecurity, such errors pose serious risks, from misleading analysts to potential adversarial exploitation. This project investigates hallucinations in three directions: (1) creating benchmarks and interpretability tools to characterize them in security contexts; (2) developing mitigation strategies such as retrieval-augmented generation, symbolic-neural hybrids, and uncertainty-aware decoding; and (3) integrating these methods into real-world workflows like vulnerability assessment, malware analysis, and penetration testing, while exploring how attackers might exploit hallucinations. Evaluation will combine accuracy metrics, human-in-the-loop studies, and red-team simulations. By bridging theory and applied system design, the work aims to advance understanding of hallucinations and improve the reliability of AI in cybersecurity, with broader implications for other high-stakes areas such as healthcare and law.

Jingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu, Jirong Yi, Myung Cho, Catherine Xu, Hui Xie, Weiyu Xu

In this paper, we study the adversarial robustness of deep neural networks (DNN) for classification against optimal classifiers. We look at the smallest magnitude of possible additive perturbations that can change a classifier's output. We provide a matrix-theoretic explanation of the adversarial fragility of DNNs for classification. In particular, our theoretical results show that the adversarial robustness of a neural network can degrade as the input dimension d increases. Analytically, we show that the adversarial robustness of neural networks can be only 1/√d of the best possible adversarial robustness of optimal classifiers. Our theories match remarkably well with empirical results. The matrix-theoretic explanation aligns with an earlier information-theoretic feature-compression-based explanation for the adversarial fragility of neural networks.

Maxime Meyer

Transformers have reshaped modern artificial intelligence, yet their theoretical foundations remain incomplete. This thesis investigates the approximation power and memory limitations of transformers. I combine tools from approximation theory and statistical learning theory to provide provable guarantees on expressivity, memorization capacity, and inherent architectural constraints. My contributions include the first rigorous proof of memory bottlenecks in prompt tuning and new results on the expressivity of transformers. The long-term goal of my doctoral research is to develop a principled theoretical framework that grounds the empirical behavior of large-scale transformer models in formal approximation-theoretic results.

Prakhar Ganesh

Model development in AI is shaped by developer decisions. While there is significant research on the opportunities and risks of multiplicity, little attention has been paid to how developer decisions impact multiplicity. My thesis focuses on (a) introducing broader frameworks to better situate and analyze developer decisions in AI, (b) identifying theoretical connections to characterize the influence of these decisions on multiplicity, and (c) operationalizing these insights across various applications, thus building responsible AI models with multiplicity.

Nitay Alon

Theory of Mind (ToM) enables agents to model others' mental states, but in mixed-motive games, this capacity can lead to deceptive behaviour and alignment risks. My research investigates how ToM affects strategic behaviour in partially observed games, contributing: (1) a formal model of ToM-driven manipulation in a preference elicitation task, (2) evidence that excessive ToM leads to paranoid-like overmentalisation, and (3) the Aleph-IPOMDP model, a framework for multi-agent systems that balances ToM reasoning with game-theoretic principles to prevent manipulation, deterring capable agents from deceiving. My work contributes to the understanding of deceptive AI, overcoming deception in multi-agent systems and applications to computational model of human cognition.

Eric Xie, Danielle Waterfield, Michael Kennedy, Aidong Zhang

Large Language Models (LLMs) have shown immense potential in education, automating tasks like quiz generation and content summarization. However, generating effective presentation slides introduces unique challenges due to the complexity of multimodal content creation and the need for precise, domain-specific information. Existing LLM-based solutions often fail to produce reliable and informative outputs, limiting their educational value. To address these limitations, we introduce SlideBot - a modular, multi-agent slide generation framework that integrates LLMs with retrieval, structured planning, and code generation. SlideBot is organized around three pillars: informativeness, ensuring deep and contextually grounded content; reliability, achieved by incorporating external sources through retrieval; and practicality, which enables customization and iterative feedback through instructor collaboration. It incorporates evidence-based instructional design principles from Cognitive Load Theory (CLT) and the Cognitive Theory of Multimedia Learning (CTML), using structured planning to manage intrinsic load and consistent visual macros to reduce extraneous load and enhance dual-channel learning. Within the system, specialized agents collaboratively retrieve information, summarize content, generate figures, and format slides using LaTeX, aligning outputs with instructor preferences through interactive refinement. Evaluations from domain experts and students in AI and biomedical education show that SlideBot consistently enhances conceptual accuracy, clarity, and instructional value. These findings demonstrate SlideBot’s potential to streamline slide preparation while ensuring accuracy, relevance, and adaptability in higher education.

Yingqi Wang, Xiaohang Luo

As generative AI rapidly enters higher education, its cognitive, motivational, and social impacts across disciplines remain underexplored. This qualitative study examines disciplinary epistemologies and digital literacy on AI-assisted academic reading among EFL Chinese students. Guided by Cognitive Load Theory and Self-Determination Theory, participants were 46 university students across Biglan's disciplinary dimensions. We analyzed 46 questionnaires and 32 interviews. Students in soft and applied fields more often report AI reducing intrinsic load, supporting deeper semantic elaboration. In pure and hard fields, students tend to use AI as an interactive tool for questioning, but multi-contextual examples are more likely to introduce extraneous load. By contrast, terminology glossing and decomposition of complex sentences are more often applied in soft and applied fields. Excessive reliance is associated with cognitive offloading and an illusory sense of mastery, shaped by digital literacy and metacognitive awareness. Socially, AI sometimes displaces routine exchanges, but when integrated into group contexts, it facilitates collaboration. The study elaborates applications of CLT and SDT by showing how disciplinary and individual factors shape AI’s cognitive and motivational roles. Practically, it proposes discipline-sensitive design principles and metacognitive prompts, pointing to deployable interventions. Ethical approval and consent were obtained.

Harshil Safi, Megha Bansal, Madhu Vadali, Barbara Bruno, Aditi Kothiyal

Theories of embodied learning emphasize that learning processes are grounded in bodily actions and interactions with the environment, suggesting that movements play a fundamental role in problem solving, decision making, and learning. This perspective holds particular relevance for making-based learning settings, where patterns of movement and spatial engagement can reveal strategic expertise. Prior research has examined distinctions between students who learned and did not learn, but manual coding of actions presents scalability and real-time application challenges. To address this gap, we develop a computer vision–based analysis pipeline for automated detection and characterization of hand movements during complex assembly tasks. In an exploratory study, we apply this approach to video data of students engaged in the assembly of a differential gearbox, quantifying metrics such as amount and speed of movement. Results indicate that learners show fewer right-hand movements than novices and exhibit reduced movement speed, with a progressive decline in speed as the task unfolds. Non-learners, by contrast, display more uneven hand movement speed. These findings, while preliminary, highlight measurable differences in actions of learners and non-learners, and therefore have potential implications for learning support. Specifically, the ability to computationally distinguish movement profiles can inform the design of adaptive learning interventions, providing real-time performance assessment and targeted feedback for making-based learning.

Alan Tsang

Multiagent systems is a key area within artificial intelligence (AI) that explores the behavior of interacting rational agents where the decisions of one agent impact others. Rooted in economic game theory, multiagent systems takes the idea of individual incentives from economic game theory and applies it to distributed computation and decentralized mechanisms. It examines not only how certain overall economic or computational goals can be accomplished, but also why individual participants will choose to cooperate with reaching that goal. While multiagent systems is grounded in rigorous mathematical theory, current pedagogical approaches often lack opportunities for students to connect abstract theory with real-world human dynamics. This disconnect is particularly pressing as AI increasingly operates in sociotechnical environments, where understanding human behavior and interaction is critical. This paper presents the first exploration of using large participation activities to facilitate experiential learning to bridge this gap. We report on a day-long resource allocation scenario involving up to 43 participants, designed to simulate multiagent interactions under pressure and with meaningful stakes, where learners can apply their theoretical knowledge to analyze and solve emerging problems. We propose ``megagames'' as a powerful pedagogical tool not only for multiagent systems, but also for other domains as well.

Khushi Malik, Amber Richardson, Tingting Zhu, Lisa Zhang

Recent work has explored the interests that draw learners to Machine Learning (ML), aiming to support their success and broaden participation in the field. However, whether strategies used in textbooks align with these interests is unexplored. We perform a thematic analysis of the introductions from ten openly available ML textbooks to identify their motivational strategies and compare them with student interests documented in prior research. We find that textbooks frequently motivate learners in their introductions by setting learning goals, previewing core ML topics to be covered, showcasing applications and current successes, and, less often, by using learner-centered strategies such as reassurance or curiosity prompts. We group these motivations into three overarching themes: theoretical, practical, and learner-centered. These motivations largely align with student interests, particularly in theory and applications, even in textbooks published before the recent surge of ML and Artificial Intelligence. These findings reveal how textbooks frame ML’s value and offer evidence-based guidance for developing future materials that better engage and support diverse learners.

Matteo Baldoni, Cristina Baroglio, Monica Bucciarelli, Sara Capecchi, Leonardo Castellani, Elena Gandolfi, Francesco Ianì, Elisa Marengo, Roberto Micalizio

Theory of mind refers to the attribution of mental states that humans ascribe to other humans or objects (such as computer-based systems). Recently, the attribution of mental states has been investigated toward Artificial Intelligence (AI) as a basic manner to capture people's engagement toward it, and people's perception about AI social skills and AI capabilities. In line with this idea, mental state attribution can be used as an indirect measure of students' understanding of AI functioning, and in particular of the kind of interactions students may have with AI systems. Too often is the case of people using generative AI systems in ways that exceed their actual ways of functioning. In our study, children of age in the range 9-12 were involved in one-shot unplugged activities concerning data and models. The unplugged activities were not aimed at teaching the theory of Machine Learning, but rather they were designed so as to provide awareness on some basic mechanisms and help developing a correct use of tools that are becoming more and more present in everyday life. This paper introduces the activities and reports the results that were achieved.

Budhitama Subagdja, Shanthoshigaa D, Ah-Hwee Tan, Iris Rawtaer

This paper introduces a novel system for in-home cognitive health assessment using ambient sensors and a machine learning technology that can robustly detect mild cognitive impairment (MCI) despite limited available data. The learned model can explain the aspects of individuals' daily lives led to the prediction, while reliably predicting MCI, providing more insights to healthcare workers for further clinical interventions. We developed the robust transparent machine learning model, based on fusion adaptive resonance theory (Fusion ART) neural network to learn individuals' daily patterns of activity from continuous sensor data in terms of a suite of digital biomarkers reflecting four key domains: physical, daily activity, cognitive engagement, and sleep patterns. Based on a longitudinal study of over one hundred participants, deployed with non-intrusive sensors in their homes to undergo parallel clinical evaluation across a period of five years, our model successfully identified individuals with MCI, achieving high predictive accuracy regardless the noisy and sparse availability of data. As a transparent neural network, the learned model can also be interpreted as classification rules to distinguish MCI from normal cognition (NC) cases based on the digital biomarkers. These results demonstrate that passively collected, sensor-derived digital biomarkers can be leveraged to indicate cognitive status and potentially providing clinically meaningful insights on the impairment conditions. We also discuss the practical challenges and lessons learned from this real-world deployment to inform future large-scale implementations of such AI-driven health monitoring systems.

Wijnand Van Woerkom, Davide Grossi, Henry Prakken, Bart Verheij

The widespread application of uninterpretable machine learning systems for sensitive purposes has spurred research into elucidating the decision-making process of these systems. These efforts have their background in many different disciplines, one of which is the field of AI & law. In particular, recent works have observed that machine learning training data can be interpreted as legal cases. Under this interpretation, the formalism developed to study case law, called the theory of precedential constraint, can be used to analyze the way in which machine learning systems draw on training data—or should draw on them—to make decisions. In the present work, we advance the theory underlying these explanation methods, by relating it to order theory and logic. This allows us to write a software implementation of the theory that can be used to compute with the definitions and give automatic proofs of the properties of the model. We use this implementation to evaluate the model on a series of datasets. Through this analysis, we characterize the types of datasets that are more, or less, suitable to be described by the theory.