论文检索

输入标题、作者或关键词,从 844 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
已筛选 KDD 2025
844篇论文
第 2 / 43 页

Haoyu Han 0001, Fali Wang, Chen Luo 0003, Hui Liu 0031, Jing Huang, Zhen Li, Zhenwei Dai, Qi He 0002, Yiwei Sun, Dawei Yin 0001 等

Large Language Models (LLMs) are revolutionizing E-Commerce by enabling product recommendation, search, classification, question answering, and advertising applications. Their increasing adoption in real-world systems underscores their potential; however, challenges persist in ensuring accuracy, efficiency, fairness, and privacy. This workshop aims to bring together researchers and industry practitioners to explore both the limitations and opportunities of LLMs in e-commerce. The workshop seeks to foster collaboration, bridge the gap between academia and industry, and drive innovation in the application of LLMs to E-Commerce through discussions on model design, algorithmic advancements, and practical deployment.

Mihajlo Grbovic, Vladan Radosavljevic, Amit Goyal, Rui Song 0006, Minmin Chen, Zhiwei Qin 0001, Katerina Iliakopoulou-Zanos, Thanasis Noulas, Hongtu Zhu, Fabrizio Silvestri

In recent years, two-sided marketplaces have emerged as viable business models in many real-world applications. In particular, we have moved from the social network paradigm to a network with two distinct types of participants representing the supply and demand of a specific good. Examples of industries include but are not limited to accommodation (Airbnb, Booking.com), video content (YouTube, Instagram, TikTok), ridesharing (Uber, Lyft), online shops (Etsy, Ebay, Facebook Marketplace), music (Spotify, Amazon), app stores (Apple App Store, Google App Store) or job sites (LinkedIn). The traditional research in most of these industries focused on satisfying the demand. OTAs would sell hotel accommodation, TV networks would broadcast their own content, or taxi companies would own their own vehicle fleet. In modern examples like Airbnb, YouTube, Instagram, or Uber, the platforms operate by outsourcing the service they provide to their users, whether they are hosts, content creators or drivers, and have to develop their models considering their needs and goals.

Yanjie Fu, Kunpeng Liu 0001, Dongjie Wang, Xiangliang Zhang 0001, Khalid K. Osman, Charu Aggarwal 0001, Suzanne M. Shontz, Huan Liu 0001, Jian Pei 0001

Machine learning traditionally emphasizes developing models for given datasets, but real-world data is often messy, making model improvement insufficient for enhancing performance. AI for data editing (AI4DE) is an emerging field that systematically improves datasets, leading to significant practical ML advancements. While experienced data scientists have manually refined datasets through trial-and-error and intuition, AI4DE approaches data enhancement as a systematic engineering discipline. AI4DE represents a shift from focusing on models to the underlying data used for training and evaluation. Despite the dominance of common model architectures and predictable scaling rules, building and using datasets remain labor-intensive and costly, lacking infrastructure and best practices. The AI4DE movement aims to develop efficient, high-productivity open data engineering tools for modern ML systems. This workshop seeks to foster an interdisciplinary AI4DE community to address practical data challenges, including data collection, generation, labeling, preprocessing, augmentation, quality evaluation, debt, and governance. By defining and shaping the AI4DE movement, this workshop aims to influence the future of AI and ML, inviting interested parties to contribute through paper submissions

Emre Eftelioglu, Naoki Abe, Ramakrishnan Kannan, Yuzhou Chen, Kathleen Buckingham, Auroop R. Ganguly, James Hodson 0003

The Fragile Earth Workshop is a recurring event in ACM's KDD Conference on research in knowledge discovery and data mining that gathers the research community to find and explore how data science can measure and progress climate and social issues, following the United Nations Sustainable Development Goals (SDGs) framework.

Haibo Ding, Shuai Wang, Kiran Ramnath, Yi Zhang, Lin Lee Cheong

Prompt engineering plays a critical role in enabling the effective use of large language models (LLMs). LLMs exhibit unpredictable sensitivity to various input factors that may result in a performance gap between prompts that are semantically indistinguishable. Therefore, LLM researchers and practitioners often optimize their prompts in an ad hoc manner due to the lack of systematic methods. Prompt optimization remains an open problem due to the rapidly evolving landscape of NLP tasks, target LLMs, and associated best practices. To address this gap, we propose the first KDD workshop on Prompt Optimization. This workshop aims to bring together researchers and practitioners working on prompt design, automatic optimization, and evaluation, fostering the exchange of ideas and methodologies. By exploring topics such as discrete and soft prompt tuning, low-resource applications, we aim to establish best practices, identify challenges, and drive future research in this critical area.

Xiquan Cui, Derek Zhiyuan Cheng, Fei Liu, Tao Ye 0001, Julian J. McAuley, Vachik S. Dave, Stephen Guo

Recommender system (RecSys) plays important roles in helping users navigate, discover, and consume massive and highly-dynamic information. Today, many RecSys solutions deployed in the real world rely on categorical user profiles and/or pre-calculated recommendation actions that stay static during a user session. However, recent trends suggest that RecSys need to model user intent in real time and constantly adapt to meet user needs at the moment or change user behavior in situ. There are three primary drivers for this emerging need of online adaptation. First, in order to meet the increasing demand for a better personalized experience, the personalization dimensions and space will grow larger and larger. It would not be feasible to pre-compute recommended actions for all personalization scenarios beyond a certain scale. Second, in many settings the system does not have user prior history to leverage. Estimating user intent in real time is the only way to personalize. As various consumer privacy laws tighten, it is foreseeable that many businesses will reduce their reliance on static user profiles. Therefore, it makes the modeling of user intent in real time an important research topic. Third, a user's intent often changes within a session and between sessions, and user behavior could shift significantly during dramatic events. Therefore, it is important to investigate more on online and adaptive recommender systems (OARS) that can adapt in real time to meet user needs and be robust against distribution shifts. Every year, the organizers survey the most important topics for OARS and propose a new workshop program. In light of the recent advancement of (multi-modal) LLMs in RecSys, in this new edition, we decide to formally add the new topic of (multi-modal) LLM models in OARS. We will invite experts and papers in the field to disseminate new knowledge and foster further advancements.

Ranak Roy Chowdhury, Yan Liu 0002, Huiming Qu, Qingsong Wen, Chen-Yu Lee, Narendra Agrawal, Alexis Roos

A supply chain is the network of entities and processes involved in the production and distribution of a commodity. Supply chains are a critical backbone across industries like retail, manufacturing, healthcare, and automotive, driving everything from product availability to operational efficiency and customer satisfaction. Modern supply chains are 1) non-cooperative, functioning as fragmented systems where isolated technologies solve individual problems without integration, and 2) unadaptable, failing to adjust to real-time data and uncertainties, like demand fluctuations and regulatory changes. As a result, frequent manual overrides are required, as even small errors can lead to significant financial losses, strained customer relationships, or reputational damage. Modern Artificial Intelligence (AI) advancements offer great potential to unify fragmented supply chains into a seamless, adaptive system while enhancing automation, decision-making, and transparency. In this workshop, we will examine critical supply chain challenges and demonstrate how AI can provide accurate, efficient, and scalable solutions.

Sudarshan Lamkhede, Moumita Bhattacharya

With the proliferation of personal computing devices and a large number of logged-in experiences, search has evolved to a stage with many different product scenarios where personalization plays a crucial role for relevance, quality, and user satisfaction. The purpose of this workshop is to have a forum where the latest research and advancements specifically on Personalization and Recommendations in Search PaRiS can be discussed in conjunction with KDD 2025. This will be the fourth instance of this workshop. We held three very successful instances of this workshop at the SIGIR 2024 PaRiS-2024, WebConf 2023 PaRiS-2023 and WSDM 2022. This year we will especially focus on applications of LLM and Generative AI to enable personalization and recommendations netflix-search in the context of search, for example, conversational assistants, while continuing to use this workshop for discussing other advances and applications in the context of personalized search and recommendations in the context of search.

Abraham Bagherjeiran, Nemanja Djuric, Kuang-chih Lee, Linsey Pang, Vladan Radosavljevic, Suju Rajan

The digital advertising field has always had challenging ML problems, learning from petabytes of data that is highly imbalanced, reactivity times in the milliseconds, and more recently compounded with the complex user's path to purchase across devices, across platforms, and even online/real-world behavior. The AdKDD workshop continues to be a forum for researchers in advertising, during and after KDD. Our website, which hosts slides and abstracts, continues to receive a large number of monthly visits and active users. In surveys during AdKDD 2019 and 2020, over 60% agreed that AdKDD is the reason they attended KDD, and over 90% indicated they would attend next year. The 2025 edition is particularly timely because of the increasing application of Graph-based NN and Generative AI models in advertising. Coupled with privacy-preserving initiatives, such as those enforced by GDPR or CCPA, the future of computational advertising is at an interesting crossroads. For this edition, we plan to solicit papers that span the spectrum of deep user understanding while remaining privacy-preserving. In addition, we will seek papers that discuss fairness in the context of advertising, to what extent does hyper-personalization work, and whether the ad industry as a whole needs to think through more effective business models such as incrementality. We have hosted several academic and industry luminaries as keynote speakers and have found our invited speaker series hosting expert practitioners to be an audience favorite. We will continue fielding a diverse set of keynote speakers and invited talks for this edition as well. As with past editions, we hope to motivate researchers in this space to think not only about the ML aspects but also to spark conversations about the societal impact of online advertising.

Aidong Zhang 0001, Vipin Kumar 0001, Yan Liu 0002

The past decade has been an inspiring time for artificial intelligence (AI) research. AI systems have transformed norms and practices across industries and have permeated the fabric of human society. Moreover, AI is ushering in a transformative technological age by making remarkable breakthroughs in a number of scientific fields such as protein structure prediction and medical imaging. There is increasing consensus in the wider scientific community that AI is poised to disrupt science by unlocking entirely new approaches, driving new scientific inquiry, and enabling greater scientific leaps with far-reaching social consequences. However, there are substantial barriers preventing science from realizing that potential, and addressing these barriers will require support for advances in AI methods and the adoption of these methods in routine scientific research. In this special day at KDD 2025, we host a series of talks by distinguished researchers on AI for science.

Filippo Maria Sposini, Aijun An

Artificial Intelligence (AI) research in Canada is driving a vibrant ecosystem of startups and industry innovation, thanks significantly to the leadership of Canada CIFAR AI Chairs. Canada AI Day at KDD 2025 will showcase cutting-edge research by some of these leading experts, with a focus on building ethical, interpretable, and accessible AI systems. The event will feature a series of invited talks followed by a panel discussion bringing together academic, government, and industry researchers to address the challenges and opportunities in developing effective and responsible AI.

Peipei Ping, Wei Ding 0003, Carl Yang 0001

The ACM KDD 2025 Health Day theme, ''Harnessing AI Opportunities in Biomedicine and Healthcare'' highlights the transformative potential of AI-driven applications in healthcare, translational biomedical research, and basic biological research. This extended abstract discusses recent advancements, challenges, and future directions, focusing on integrating AI-ready data sets, interdisciplinary collaborations, and ethical AI practices. It aims to catalyze discussions on the potential of AI ecosystems in revolutionizing healthcare and related fields.

Jun Huan, Xiangyu Zhang, Ye Xing, Wee Hyong Tok, Ruzica Piskac

Generative AI and the use of large language models (LLMs) are changing the way we work, create, play, and live. As we have witnessed in the past few years, there is significant progress in training LLMs to have a deep understanding of the semantics of language so that such models begin to perform ''reasoning''. (Human) Reasoning is the process of applying logic to derive conclusions based on new or existing information with the goal of finding the truth. Reasoning is a form of high-level human intelligence. There are many types of reasoning: mathematical reasoning, common sense reasoning, temporal reasoning, among others. Multi-hope reasoning with LLM is an emerging capability for LLMs with tens of billions of parameters. Such ''reasoning models'', including Sonnet 3.7, Chat GPT O1, have powered important application areas such as AI4coding, agentic workflow, among others. The first KDD AI Reasoning Day is a special event that we organize in order to increase the awareness of this important research topic for the research community. We bring leaders from industry and academia to present the latest progresses on improving LLM's reasoning capability and enabling reasoning for different application development.

Ebrahim Bagheri, Faezeh Ensan, Calvin Hillis, Reihaneh Rabbany, Robin Cohen, Benjamin C. M. Fung, Sébastien Gambs

This special day event on Responsible Artificial Intelligence (AI) brings together researchers, practitioners, and policymakers to explore how data mining and machine learning systems can be designed to align with ethical principles, societal values, and human well-being. As AI technologies increasingly influence decisions in healthcare, finance, governance, and social systems, there is a critical need to develop frameworks that embed fairness, accountability, and privacy directly into the foundations of knowledge discovery. This full-day event will feature a mix of invited talks, interactive debates, expert panels, and peer-reviewed research presentations, all focused on the practical integration of ethical design into data-driven systems. The Responsible AI Day builds on the success of Canada's NSERC CREATE Program on Responsible AI, an interdisciplinary initiative training the next generation of AI researchers across computer science, law, bioethics, public health, and media studies. Topics will span scalable AI governance, privacy-preserving computation, algorithmic bias mitigation, and the socio-legal tensions emerging in generative AI. By positioning responsible AI as a sociotechnical challenge, this special day aligns with KDD's mission of advancing data science that is not only technically robust but also socially conscious.

Eric Xu, Lin Deng 0001

In an era where justice and accountability increasingly depend on digital evidence, Large Language Models (LLMs) offer transformative potential for digital forensics. This three-hour Hands-on tutorial explores how LLMs can automate investigations, reveal hidden insights, and enhance evidence analysis. Through real-world case studies, interactive exercises, and hands-on labs, participants will learn to leverage LLMs for tasks such as entity identification, evidence processing, and knowledge graph reconstruction. Designed for professionals, researchers, and students, this collaborative learning experience equips attendees with practical skills to innovate in digital forensics. As LLMs reshape the field, this tutorial underscores their role in improving justice outcomes, strengthening accountability, and advancing the future of digital investigations.

Zhaopeng Qiu, Jingqi Zhang, Shuang Yu, Shuai Zhang, Junjie Lai

With reasoning models like DeepSeek-R1 and OpenAI's o1 demonstrating breakthrough capabilities in complex problem-solving, there is growing interest in the AI community about how to unlock similar capabilities in other large language models (LLMs). This hands-on tutorial dives into practical methods for building reasoning capabilities in LLMs through two primary approaches: knowledge distillation from advanced reasoning models and post-training with reinforcement learning techniques. Participants will learn how to transfer reasoning capabilities from cutting-edge models like DeepSeek-R1 into smaller LLMs such as Qwen and Llama, and then explore how reinforcement learning can take these capabilities even further. Through interactive Jupyter notebook, the participants will exercise through the entire process. By the end of this session, participants will be equipped with practical knowledge in how to incentivize reasoning capabilities into LLMs, understand how to use various frameworks for this task, and leave with hands-on experience that can be applied to their own projects. The related materials are available at https://zpqiu.github.io/reasoning-model-tutorial-kdd2025.

Eliana Pastor, Eleonora Poeta, André Panisson, Alan Perotti, Gabriele Ciravegna

As deep learning systems become pervasive, the demand for trustworthy and transparent AI continues to grow. Traditional feature attribution methods, however, often lack robustness and alignment with human reasoning. This tutorial moves beyond feature attribution by introducing participants to two complementary interpretability paradigms: Concept-Based Explainable AI (C-XAI) and Mechanistic Interpretability. C-XAI provides explanations grounded in high-level, human-interpretable concepts, bridging the gap between model reasoning and human understanding. In parallel, mechanistic interpretability--a quickly emerging field--focuses on reverse-engineering neural networks to uncover and disentangle the internal mechanisms that give rise to human-understandable representations. Through interactive coding sessions and hands-on exercises, attendees will gain practical experience implementing, evaluating, and comparing a variety of C-XAI and mechanistic interpretability techniques. By the end of the tutorial, participants will be equipped with a modern interpretability toolbox and a deeper understanding of how to apply them in real-world scenarios.

Quentin Nater, Mourad Khayati, Philippe Cudré-Mauroux

Although missing gaps are common in time series data, most existing imputation libraries have a narrow focus. They typically rely on a limited set of techniques and make overly simplistic assumptions about the nature of missing data. Consequently, they fail to model the true intricate complexity of real-world time series. To overcome these challenges, we developed ImputeGAP, a versatile and comprehensive library for time series imputation. ImputeGAP supports a wide range of imputation algorithms and modular missing data simulation, catering to datasets with varying characteristics. It also streamlines imputation analysis with features such as automated hyperparameter tuning, benchmarking, explainability, and downstream evaluation. In this tutorial, we will provide an engaging hands-on tutorial where you will learn time series imputation using the powerful Python library, ImputeGAP. The session is divided into two parts. In the first part, we will dive into building an end-to-end imputation workflow with the library, where you will explore real-world missingness patterns simulation, leverage automated tuning for optimal imputation, and benchmark imputation techniques-all with extensive customization options. In the second part, we will unlock advanced functionalities, including assessing the impact of imputation on downstream analytics and understanding how time series features influence imputation outcomes. Whether you are a researcher or practitioner, this tutorial will provide you with the expertise to handle missing data in time series. The ImputeGAP library is accessible at: https://imputegap.readthedocs.io.

Matthew B. A. McDermott, Justin Xu, Teya S. Bergamaschi, Hyewon Jeong, Simon A. Lee, Nassim Oufattole, Patrick Rockenschaub, Kamile Stankeviciute, Ethan Steinberg, Jimeng Sun 0001 等

Health AI suffers from a systemic reproducibility crisis that irreparably hinders research in this space across academia and industry. To combat this and empower researchers in the health AI space, we propose a comprehensive interactive tutorial introducing the ''Medical Event Data Standard'' (MEDS) and its growing open-source ecosystem. Working in MEDS allows you to more easily build AI models over public or private longitudinal EHR datasets and to readily benchmark existing, published models against contributions on local datasets and tasks. MEDS simplifies the construction of AI models on longitudinal Electronic Health Record (EHR) datasets and enables straightforward benchmarking against established models. Reflecting its growing adoption, MEDS is utilized at over 15 institutions across 8 countries, features 7+ open-source tools, supports 10+ published models, and provides publicly available Extract-Transform-Load (ETL) pipelines for major public EHR datasets. A KDD tutorial offering practical experience with MEDS will significantly enhance reproducibility and comparability in health AI research. In this tutorial, we will teach attendees how to (1) transform datasets into the MEDS format(2) pre-process MEDS data for modeling needs(3) build highly effective, efficient, AI models for diverse predictive tasks on their datasets, and (4) contribute their results to MEDS-DEV, a decentralized benchmark enabling robust evaluation against meaningful baselines. Participants will engage in collaborative, minimal-dependency Jupyter notebook exercises, guided through each step by structured instruction and practical coding sessions. Attendees will leave equipped with practical knowledge to build reproducible, state-of-the-art AI models within the MEDS ecosystem.

Yozen Liu, Tong Zhao 0003, Matthew Kolodner, Kyle Montemayor, Shubham Vij, Neil Shah

Recent advances in graph machine learning (GML) and Graph Neu- ral Networks (GNNs) have sparked significant practical interest given the ability to model complex relationships between entities. Despite rapid progress in GNN designs, scalability remains a major challenge. Industry applications require solutions that can handle graphs with billions of nodes and edges efficiently. GiGL (Gigantic Graph Learning) is an open-source library from Snapchat, designed for large-scale distributed training and inference with GNNs. It seamlessly integrates with popular open-source GNN libraries like PyTorch Geometric (PyG). GiGL provides simplified configurable interfaces with minimal modeling code requirements, providing in- dustrial practitioners a straightforward way to apply GNNs to large- scale applications and enabling academics to conduct large-scale experiments. At the same time, it enables complex modeling capabil- ities desirable for modeling iteration. In this hands-on tutorial, we will demonstrate how GiGL addresses the scalability challenge in GNNs and provide a step-by-step guide for attendees to complete end-to-end training and inference with GiGL on industry-scale graphs. By the end of our tutorial, participants will have hands-on experience in training GNNs on graphs with billions of nodes and edges - capabilities not easily achievable with open-source graph learning libraries like PyG alone. We anticipate strong interest and participation from both industrial practitioners working on GNN applications and academics conducting large-scale experiments.