论文检索

输入标题、作者或关键词,从 11,272 篇学术成果中精准定位

会议来源 已选 1 项

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

已选择 1 个会议
支持跨会议组合检索,PDF 均跳转至官方来源
已筛选 AAAI
11,272篇论文
第 249 / 564 页

Ashlesha Akella, Krishnasuri Narayanam

Ensuring data quality in large tabular datasets is a critical challenge, typically addressed through data wrangling tasks. Traditional statistical methods, though efficient, cannot often understand the semantic context and deep learning approaches are resource-intensive, requiring task and dataset-specific training. We present an automated system that utilizes large language models to generate executable code for tasks like missing value imputation, error detection, and error correction. Our system aims to identify inherent patterns in the data while leveraging external knowledge, effectively addressing both memory-dependent and memory-independent tasks.

Zijian Zhang

My research direction is about the pancreatic cancer diagnosis system. Pancreatic cancer, as one of the cancers with the highest mortality rate, has always been a difficult problem in world medicine. I hope that through my efforts, I can contribute to the integration of AI and medicine, and contribute to increasing the probability of early diagnosis of pancreatic cancer.

Katherine Xu

Falls among older adults pose a significant public health challenge, impacting quality of life and healthcare costs. This research proposal aims to develop an innovative AI-driven personalized fall prevention system for older adults, leveraging advanced machine learning techniques in computer vision, natural language processing, and reinforcement learning. The proposed system will encompass five key components: (1) Advanced pose estimation and activity recognition using HRNet with attention mechanisms and hybrid LSTM-GCN models; (2) Personalized risk assessment through multi-modal deep learning, combining CNNs, RNNs, and federated learning for privacy-preserving distributed training; (3) Adaptive intervention strategies employing Deep Q-Networks and model-based reinforcement learning with GAN-simulated environments; (4) Human-AI interaction utilizing SHAP values for explainable AI and fine-tuned GPT-3 for natural language communication; and (5) Privacy-preserving techniques including differential privacy and homomorphic encryption. The research will be conducted over a five-year period, involving data collection, model development, large-scale testing, and clinical trials. Expected outcomes include a scalable, privacy-preserving AI system capable of significantly reducing fall incidents among older adults, thereby improving quality of life and reducing healthcare costs. This interdisciplinary research contributes to advancing AI techniques in real-world healthcare applications while addressing critical ethical and privacy concerns, potentially transforming elderly care on a global scale.

Devin Wingfield

Artificial intelligence (AI) has improved significantly in recent decades, and, along with it, its applications to real-world scenarios. AI has been used within a wide variety of fields like health care and e-commerce, however, AI has yet to integrate with the agriculture industry. With the help of machine learning, AI can begin to integrate with the industry via a research assistant. The model will assist researchers conduct experiments by giving treatment methods that are best suited for the experiment rather than relying on the expertise of the researcher. This will help research within the industry to become more efficient and less error prone. To accomplish this, the model will use a Knowledge Graph created by the IDIR lab that converts the large CSV files into a graph that can be queried and then summarized by the model.

Anna Serbina

Autonomous robots are essential for navigating and collecting data in hazardous environments where human intervention is impractical. Current methods often result in inefficiencies, missed high-quality imagery, and inadequate coverage in critical areas such as environmental monitoring, disaster response, and medical diagnostics. The absence of intelligent viewpoint selection leads to redundant data and poor image quality, limiting robotic effectiveness. This research proposes a framework that utilizes reinforcement learning and information-theoretic approaches to optimize viewpoint selection, aiming to enhance data collection efficiency and image quality while ensuring safety. This work has the potential to transform industries reliant on precise visual data and significantly improve medical robotics, enabling better diagnostics and patient care.

Mariana Risco Cosavalente

This project leverages Visual Question Answering (VQA) to promote Peruvian gastronomy by utilizing a culturally rich dataset and advanced models such as LLaVA-1.5 and GPT-2 Large. The evaluation will comprise both automated metrics and culinary expert assessments. This system addresses regional variations in dish names, promotes inclusivity by involving Peruvians from diverse regions in dataset construction, and enhances cultural representation.

Sophia Pi

Biological neural systems often represent information on low-dimensional manifolds that reflect the topology of their encoded variables. This suggests that neural activity can be naturally organized in geometrically meaningful ways, as seen in rodent head direction cells forming circular manifolds. This proposal examines whether artificial neural networks (ANNs) trained on tasks with well-defined topologies—such as planar or spherical coordinates from autonomous driving datasets like Apolloscape, cyclic temporal variables, or graph-structured road networks—develop similar low-dimensional representations aligned with the variables' inherent topology. We consider convolutional and vision transformer models for image data, graph neural networks for road network graphs, and 3D or point-based models for LIDAR point clouds, analyzing their internal activations with dimensionality reduction and topological data analysis. If successful, this approach not only elucidates the nature of internal representations in ANNs but also offers insights into the computational principles that bridge artificial systems and biological cognition.

Thao Pham

Many existing benchmarks, such as MMLU, are limited to measuring large language models’ (LLM) true task understanding due to their reliance on statistical patterns in the training data. We suggest new approaches to improve how benchmarks can capture task-specific understanding in LLMs, revealing insights into their reasoning ability.

Chidiogo Nwabuike

This paper explores the application of computer vision technology as a proactive solution to prevent Road Traffic Accidents in Nigeria. By leveraging machine learning algorithms and real-time video analysis, computer vision can reduce incidents caused by human error. The research focuses on designing an autonomous but elaborate system that monitors traffic patterns, road irregularities and triggers automated interventions when risky conditions are detected. The aim is to suggest that computer vision can be pivotal in enhancing road safety and reducing traffic-related fatalities in Nigeria.

Gerald Ketu Ndawula

Healthcare diagnostics, especially in underserved communities, faces critical gaps in accessibility and accuracy. African Americans experience significant disparities in mental health care, often receiving delayed or inadequate treatment. This research proposes a diagnostic copilot, an AI-powered assistant designed to work alongside healthcare professionals. Using Knowledge-Infused Learning (KIL) and multi-turn conversations, the system integrates clinical knowledge and patient input to deliver actionable, explainable diagnoses in real-time. By engaging with both patients and clinicians, the copilot aims to reduce disparities, enhance trust, and improve diagnostic accuracy in mental health care.

Jessica E. Liang

Diffusion Models (DMs) offer robust tools for addressing uncertainty and enhancing adaptability in robotics. This work explores their application to trajectory generation, 3D image synthesis, and interpretable scene understanding. For trajectory planning, we propose using colored Gaussian noise to improve robustness and temporal coherence. In 3D image generation, Transfer Entropy enhances information flow between textual and visual modalities for more coherent outputs. Partial Information Decomposition (PID) is leveraged to improve model interpretability and efficiency in scene generation. Rigorous evaluation will assess trajectory quality, robustness, and real-world transferability, aiming to advance autonomous decision-making and scene understanding in robotics.

Jiarui Li

When digitizing documents using conventional equipment, shadows often appear, posing significant challenges to the visual quality and readability of the digital copies. Given that the removal of document shadows typically involves complex image processing and computational tasks, which require substantial computational resources and time, the cost can become prohibitive, limiting the practicality and efficiency of shadow removal algorithms. This research aims to address the critical task of designing a model capable of achieving superior shadow removal effects. We propose a deep learning model for document shadow removal that harnesses Sobel text prior and ground truth masks as supervision. This prior knowledge encapsulates regular information regarding document structure and shadow formation, thereby enhancing its ability to utilize edge information for shadow removal. Additionally, the integration of prior knowledge and supervised learning can help the model learn more quickly, reducing the amount of information the model needs to process and improving its efficiency.

Gautam Jajoo

In recent years, federated learning (FL) has emerged as a promising technique to enable decentralized training of models without the need for data centralization, addressing privacy concerns and reducing communication over-head. The challenge, however, lies in scaling federated systems to accommodate clients with different computational capabilities. The heterogeneity of clients in terms of data, model structures, and computational resources presents significant challenges. Addressing these challenges can lead to more robust and efficient FL systems, making it possible to leverage diverse data sources and computational environments. Here we propose a system where small language models run on heterogeneous cli-ents while a large, more powerful model at the server aggregates their contributions. This architecture leverages the strengths of both small, task-specific client models and a large server model to enhance generalization and efficiency. This is important because it addresses the growing need for scalable, privacy-preserving systems that can operate in diverse environments with varying resources. Through such a system we intend to contribute to the AI field by improving the efficiency of federated learning systems while enhancing their adaptability to real-world applications.

Iteoluwa Ibitoye

The global expansion of Artificial Intelligence (AI) has highlighted significant challenges in inclusivity and representation, particularly for underrepresented communities. Current AI systems often fail to accommodate diverse linguistic and cultural contexts, resulting in biases in name pronunciation, language preservation, and communication. This research proposes a framework for advancing inclusivity in AI through Natural Language Processing (NLP) and Reinforcement Learning (RL). The envisioned system could integrate with home assistants like Siri and Alexa, enabling real-time interactions in local languages while maintaining cultural relevance. Key proposed features include accurate pronunciation of names, conversational capabilities in underrepresented languages, and an interactive platform where users can learn their language, history, and cultural heritage. By leveraging transformer-based models and adaptive RL frameworks, this research aims to explore solutions that bridge the gap in AI inclusivity for low-resource languages and culturally diverse populations.

Kennedy Hecker

This paper outlines a proposal regarding the use of machine learning, specifically a long-short term model, to increase the military’s effectiveness and safety protocols. The approach is to collect data from weapons training and apply it to a model that can distinguish between weapon activities. By training the model on a dataset that consists of several common weapons activities, we hope to improve commanders' understanding of their troop's performance and readiness. The evaluation will consist of examining the loss of the model, its accuracy, and analyzing activities it frequently confused. This work will extend the current research in soldier activity recognition by introducing weapon activity recognition.

Precious Donkor

This paper investigates implicit biases in large language models (LLMs) triggered by subtle contextual cues. Through experiments, the study examines how these biases influence model outputs in domains such as healthcare and hiring. A framework for mitigating stereotype reinforcement is proposed, along with strategies to refine prompts and reduce biased responses. The goal is to improve fairness in AI-driven applications by addressing these biases and enhancing model equity.

Soyon Choi

Following the rapid rise of deep learning (DL) and generative artificial intelligence (GenAI), it is imperative that we gain a better understanding of how these machine learning (ML) systems actually learn. What information are DL models retaining from the training data? What reasoning capabilities do these models have? In my proposed project, I aim to tackle these pressing questions through use of an adversarial lens.

James Blossom Eleojo

Leaf based diseases in tomatoes such as early blight, late blight, and septoria leaf spot, pose a significant threat to global food security and have substantial economic impacts. Early detection of these diseases is crucial for improving crop yields. This paper explores the use of vision-language models (VLMs) for detecting tomato leaf diseases by fine-tuning a pre-trained model on a large dataset of tomato leaf images with corresponding disease annotations. This approach enhances disease detection accuracy and enables multi-modal learning, real-time monitoring, and automated diagnosis, offering promising applications in precision farming and food production.

Hunnain Arsalan

This research proposes an AI-driven early warning system to predict patient deterioration in real-time using electronic health records (EHRs) and wearable devices. Leveraging deep learning techniques, such as recurrent neural networks (RNNs) for sequential data and convolutional neural networks (CNNs) for pattern recognition, the system adapts dynamically through reinforcement learning. Evaluation strategies include retrospective and prospective studies in clinical settings, measuring prediction accuracy and impact on patient outcomes. If successful, this system has the potential to save lives, reduce ICU admissions, and transform healthcare into a proactive, data-driven field.

Nicholas Abram

The lack of personalization in early education can often leave students with weak foundational skills, causing said students to be behind in their studies. Personalized learning, the idea of tailoring a unique lesson plan to a student, has been shown to improve the understanding of content learned. Robots utilizing personalization techniques in educational settings, coined social robots, have been able to form a connection with students, thereby keeping them engaged while learning. This proposal seeks to study the effects of AI-driven social robotic tutors coupled with personalized learning on early childhood education. The study will consist of five groups of K-4 students: two groups learning while utilizing both a social robot and a tablet (one group with personalized learning and the other without), the two groups interacting with only the tablet (with and without personalization), and the last group learning utilizing both a non-personalized learning tablet and a non-social robot. This study aims to determine whether the combination of robotic interaction and personalized learning leads to better outcomes than solely tablet-based or non-personalized methods. This study will focus on teaching mathematics to the participants. Pre and post-tests will measure learning progress, and the influence of robot interaction on student engagement will also be evaluated. It is expected that the students with social robotic tutors and personalized learning tablets will show the greatest knowledge retention, outperforming all other categories. These findings could have significant implications for the integration of AI and robotics in early education, potentially revolutionizing how personalized learning is implemented therefore improving educational outcomes for young learners.