The neural tangent kernel (NTK) has emerged as an important tool in recent years, both for developing a theoretical understanding of deep learning as well as for various applications. Even though recursive closed form expressions have been derived for computing the NTK, these become computationally expensive as the complexity of a network increases. Recent papers have looked at reducing this complexity using various sketching techniques along with random features. Building on these techniques, we propose an additional optimization step which results in better approximation of the NTK.
论文检索
输入标题、作者或关键词,从 11,272 篇学术成果中精准定位
A vast amount of textual data is added to the internet daily, making utilization and interpretation of textual data difficult and cumbersome. As a result, automatic text summarization is crucial for extracting relevant information, saving precious time. Although many transformer models excel in summarization, they are constrained by their input size, preventing them from processing texts longer than their context size. This study introduces several novel algorithms that allow any LLM to efficiently overcome its input size limitation, effectively utilizing its full potential without any architectural modifications. We test our algorithms on texts with more than 70,000 words, and our experiments show a significant increase in BERTScore with competitive ROUGE scores.
scMBERT: A Pre-Trained Deep Learning Model for Single-Cell Multiomic Data Representation and Prediction (Student Abstract)
PDF ↗Recent advancements in single-cell sequencing technologies enable the measurement of multiple modalities in individual cells, offering insights into the transcriptome and regulome in various biological systems and human diseases in an unprecedented resolution. However, effectively using these ultra-high-dimensional and large-scale multiomic data to understand gene regulation remains challenging. Inspired by the success of adapting large language models into the genomics field, we develop scMBERT, a BERT framework-based pre-trained deep learning model using single-cell multiomic data. We showed that scMBERT increases model flexibility and performance in downstream tasks like cell type annotation and batch-effect correction, demonstrating the potential of leveraging multiomic data to improve single-cell genomic data analyses.
Communication Accommodation Between Large Language Models and Users Across Cultures (Student Abstract)
PDF ↗The increasing adoption of conversational agents powered by large language models (LLMs) raises questions about its effects across culturally diverse interactions. While these agents are linguistically versatile and multilingual, their ability to adapt along cultural dimensions--defined as geographically and communally nurtured sets of values and behavioral norms--lacks close scrutiny of both their design and deployment. To achieve inclusive conversational AI, it is essential to understand how agents adapt to users from diverse cultural backgrounds. In this study, we analyze dialogues between human users from different countries and LLM-powered agents to examine how both parties adapt their word use, a salient aspect of linguistic styles, toward one another throughout casual conversations. Our analysis reveals that LLMs exhibit varying degrees of style matching based on users' national cultures and demonstrate asymmetric adaptation when interacting with culturally diverse users. Moreover, we observe a reciprocal dynamic where both the LLMs and users from certain cultures adjust their styles in response to one another. Additionally, our findings support the hypothesis that LLMs and users naturally converge in conversational styles over the course of interactions, mirroring the dynamics of human conversations that accommodate and converge. To develop localized and culturally aware agents, there's a potential to utilize such cross-cultural convergence process during fine-tuning to align LLMs.
Large language models (LLMs) now turn their attention to search. Recently, Thought of Search (ToS) proposed defining the search space with code, having an LLM produce that code. ToS requires a human in the loop, collaboratively producing a sound successor function and goal test, achieving impressive 100% accuracy on all the tested datasets. In this work, we automate ToS (AutoToS), completely taking the human out of the loop of solving planning problems. AutoToS guides the language model step by step towards the generation of sound and complete search components, through feedback from both generic and domain specific unit tests. We achieve 100% accuracy, with minimal feedback iterations, using LLMs of various sizes on all evaluated domains.
The main goal of Few-Shot learning algorithms is to enable learning from small amounts of data. One of the most popular and elegant Few-Shot learning approaches is Model-Agnostic Meta-Learning (MAML). In this paper, we propose a novel framework for Bayesian MAML called BH-MAML, which employs Hypernetworks for weight updates. It learns the universal weights point-wise, but a probabilistic structure is added when adapted for specific tasks. In such a framework, we can use simple Gaussian distributions or more complicated posteriors induced by Continuous Normalizing Flows.
We propose the Next Sentence Prediction (NSP) task as a simple, objective, scalable, automated way to test ChatGPT’s text comprehension. Given a context excerpted from a children’s story, the task is to distinguish the next story sentence from a later sentence in the story. We analyze how ChatGPT’s performance on this task is related to various features of the text, using data from English and Swahili children’s stories.
Synthesizing electronic health records (EHR) is essential for addressing data scarcity, bias, and fairness in healthcare models. EHR data are inherently multimodal and sequential, encompassing structured codes, clinical notes, medical images, and irregular time intervals. Traditional generative models like GANs and VAEs struggle to capture these complexities, while diffusion-based models offer improvements but remain limited to task-specific applications. To address these challenges, two diffusion-based models, MedDiffusion and EHRPD, have been developed. MedDiffusion enhances health risk prediction by generating synthetic patient data and capturing visit-level relationships, while EHRPD generates sequential, multimodal EHR data, incorporating temporal interval estimation to improve diversity and fidelity. Future work aims to overcome limitations in multimodal data generation by developing a generalized model capable of handling diverse modalities simultaneously, expanding the applicability of EHR data generation across healthcare tasks.
While advances in machine learning and the expansion of massive datasets have significantly improved predictive accuracy, the translation of these predictions into actionable decisions—alongside a robust understanding of associated risks—remains underexplored. My research focuses on developing methodology and theory in data-driven decision-making and uncertainty quantification that effectively address core data challenges. This paper presents two connected pillars of my research: data-driven contextual optimization, uncertainty quantification and reduction.
Economic growth and development require a consistent supply of energy. Energy which has mainly been supplied from fossil fuels. The impacts of these on the environment such as global warming, raised an alarm on their use. As a result other sources of energy such as wind energy are used as alternatives for electricity production. Wind energy assessment nevertheless faces barriers due to its stochastic nature. This later creates various regimes, which traditional models can't always fit thereby producing poor estimates. In this work, we aim to use Large Language Models (LLMs) to predict the wind potential in a given location. Through this approach, we aim at lifting the barrier on energy problems in developing countries by providing knowledge on the state of wind energy in given locations.
Advancing Intelligent Software Development and Trustworthy Models Through the Synergy of Software Engineering and LLMs
PDF ↗Integrating Large Language Models (LLMs) into software engineering unlocks new opportunities to automate manual processes but raises challenges around reliability, safety, and scalability. My research centers on this synergy, with two key objectives: first, harnessing LLMs to solve software engineering tasks traditionally dependent on labor-intensive, domain-specific methods, and second, applying robust software engineering principles to improve LLM safety and performance. This dual focus creates a powerful feedback loop, where LLMs drive innovation while engineering rigor ensures these systems meet the high standards required for real-world applications.
While machine learning (ML) models of today have the potential to be useful in many societal applications, they also harbor the potential for great harm, be it perpetuating biases or compromising privacy. To prevent these harms, many (evolving) regulatory guardrails have been put in place; for instance European Union's GDPR and Biden's Executive Order which demand explainability, privacy, fairness and so on from models deployed in societal applications. Yet, most technical solutions in the Trustworthy ML literature which claim to meet these regulatory requirements are brittle and often fail at the task in hand. To this end, my research aims to make the field of Trustworthy ML reliable using mainstay concepts of Measurement, Mitigation and Maintenance. With these concepts, I develop end-to-end solutions for trustworthy ML by (1) exploring the limitations of existing approaches and (2) providing principled novel solutions exploiting interconnections with cryptography.
Time sequences are essential in fields such as finance, healthcare, and environmental science, where understanding temporal dependencies and making accurate predictions are crucial. These sequences often exhibit complexities like nonlinearity, noise, and concept drift. Traditional models struggle to capture the intricate dynamics of multivariate and co-evolving sequences, particularly in contexts where relationships between variables shift unpredictably. This thesis introduces a range of Kernel Representation Learning (KRL) methodologies to address these challenges. We develop kernel self-representation learning to capture the temporal dependencies and hidden structures, while identifying concept drift in co-evolving sequences. Additionally, we explore theoretical connections between KRL and advanced deep-learning models. The proposed methods are validated through real-world applications, showing improvements in predictive accuracy, interpretability, and robustness.
Language Model Meets Prototypes: Towards Interpretable Text Classification Models through Prototypical Networks
PDF ↗Pretrained transformer-based Language Models (LMs) are well-known for their ability to achieve significant improvement on NLP tasks, but their black-box nature, which leads to a lack of interpretability, has been a major concern. My dissertation focuses on developing intrinsically interpretable models when using LMs as encoders while maintaining their superior performance via prototypical networks. I initiated my research by investigating enhancements in performance for interpretable models of sarcasm detection. My proposed approach focuses on capturing sentiment incongruity to enhance accuracy while offering instance-based explanations for the classification decisions. Later, we develop a novel white-box multi-head graph attention-based prototypical framework designed to explain the decisions of text classification models without sacrificing the accuracy of the original black-box LMs. In addition, I am working on extending the attention-based prototypical framework with contrastive learning to redesign an interpretable graph neural network for document classification, aiming to enhance both the interpretability and performance of the model in document classification.
Foundation models in general domains have leveraged multimodal knowledge graphs to great effect, yet the healthcare sector lacks such comprehensive structures, presenting a significant gap in current research. Based on previous exploration with pure data-driven approaches, this proposal describes a two-stage project aiming to enhance multimodal healthcare foundation model with domain knowledge. The first stage is to construct a robust multimodal healthcare knowledge graph based on established healthcare taxonomies, such as UMLS, and enriched with data from multimodal clinical databases like MIMIC-CXR. This knowledge graph will incorporate medical images as cross-modal instances linked to healthcare terminologies, enhancing the depth and applicability of the graph. In the second stage, the knowledge graph will serve as a foundational tool in training healthcare foundation models with enhanced capabilities, particularly in reducing hallucination and managing concept ambiguity through the novel use of reinforcement learning techniques like Direct Preference Optimization (DPO). This research is expected to make significant contributions to the domain of healthcare AI by enabling more accurate, reliable, and explainable AI-driven diagnostics and interventions.
Design and Evaluation of a Generative Artificial Intelligence-Based Tool for Students with Learning Disabilities in Kenya
PDF ↗Students with learning disabilities (LDs) face significant challenges in key academic areas such as reading comprehension, cognitive organization, self-expression, mathematics, and handwriting. These difficulties increase their susceptibility to discrimination and mental health related issues. Although existing studies have primarily focused on AI’s diagnostic capabilities, there is limited research examining how Generative AI (GenAI) can be utilized to produce measurable learning outcomes and enhance learning experiences for students with LDs. Moreover, GenAI is increasingly gaining prominence in educational settings. Therefore, the relationship between GenAI tools, LDs, and instructional methods needs to be further examined. This research aims to develop a comprehensive theoretical framework for helping design and implement tools specifically tailored to the unique needs of students with LDs. A prototype based on this framework will be implemented in selected educational settings to assess its effectiveness in improving learning outcomes and providing targeted support to students with LDs. The prototype will provide mobile phone integration to ensure scalability and enhance educational accessibility. The expected findings will contribute to the promotion of more inclusive learning environments for students with LDs.
Large Language Models (LLMs) have shown promise in educational applications, but challenges such as hallucinations, lack of contextual relevance, and limited personalization impede their practical adoption. To address these issues, my research introduces MerryQuery, an LLM-powered educational agent that integrates Retrieval-Augmented Generation (RAG), rule-based content control, and Reinforcement Learning from Human Feedback (RLHF). The system features a dynamic learning profile module for adaptive personalization and a multi-step verification framework that cross-checks responses against external sources to enhance trustworthiness. A functional prototype of MerryQuery is being piloted in a real-world classroom. Preliminary results demonstrate improved response reliability and student understanding.
Deploying machine learning (ML) models in high-stakes domains such as healthcare and autonomous systems requires reliable uncertainty quantification (UQ) to ensure safe and accurate decision-making. Conformal prediction (CP) offers a robust, distribution-agnostic framework for UQ, providing valid prediction sets that guarantee a specified coverage probability. However, existing CP methods are often limited by assumptions that are violated in real-world scenarios, such as non-i.i.d. data, and by a lack of integration with modern machine learning workflows, particularly in large generative models. This research aims to address these limitations by advancing CP techniques to operate effectively in non-i.i.d. settings, improving predictive efficiency without sacrificing theoretical guarantees, and integrating CP directly into model training processes. These developments will enhance the practical applicability of CP for a wide range of ML tasks, enabling more reliable and interpretable models in high-stakes applications.
Towards Autonomous Network Management: AI-Driven Framework for Intelligent Log Analysis, Troubleshooting and Documentation
PDF ↗As modern network management grows increasingly complex, administrators are tasked with navigating vast volumes of log data, often resulting in inefficiencies, errors, and operational challenges. My doctoral research addresses these pressing issues by leveraging advanced AI techniques to minimize human intervention and pave the way for fully automated network operations. I propose a novel AI-driven framework that integrates Large Language Models (LLMs) with Retrieval-Augmented Generation (RAG) and a human-in-the-loop process to effectively automate key network management tasks, including log analysis, troubleshooting recommendations, and documentation generation. By enhancing the accuracy and efficiency of these tasks, this study aims to improve network reliability, reduce operational complexity, and contribute to the evolution of self-running networks.
Large Language Models (LLMs) and Generative AI (GenAI) have markedly changed the landscape of many fields, including education. While these tools have significant capabilities, they also require understanding to effectively and responsibly use them. Additionally, little work has been done to evaluate how these tools can best benefit education at the secondary level, with design insights from instructors. My work focuses on informing secondary instructors of these tools, receiving their input on how to make these tools work best for them, and finally using this input to create and evaluate an in-class Retrieval-Augmented Generation (RAG)-based chatbot for their students to use to improve learning outcomes. This work aims to bridge the gap between the latest in computing technology and secondary education classrooms.