How might reinforcement-learning based covert social influence operations (CSIOs) be run, given that the CSIO agent wants to maximize influence and minimize discoverability of malicious accounts? And how successful can they be, given that both social platform bot detectors and humans might report them to the social platform? To answer these questions, we propose RL_CSIO, a methodology based on reinforcement learning (RL) for running CSIOs. We ran 4 CSIOs with IRB-approval over a period of 5 days using a panel of 225 human subjects. We explore 8 research questions based on the data collected. The results show that RL_CSIO agents successfully trade off influence and discoverability - but in ways that are nuanced and unexpected.
论文检索
输入标题、作者或关键词,从 2,072 篇学术成果中精准定位
Recently, research on Text-Attributed Graphs (TAGs) has gained significant attention due to the prevalence of free-text node features in real-world applications and the advancements in Large Language Models (LLMs) that bolster TAG methodologies. However, current TAG approaches face two primary challenges: (i) Heavy reliance on label information and (ii) Limited cross-domain zero/few-shot transferability. These issues constrain the scaling of both data and model size, owing to high labor costs and scaling laws, complicating the development of graph foundation models with strong transferability. In this work, we propose the GraphCLIP framework to address these challenges by learning graph foundation models with strong cross-domain zero/few-shot transferability through a self-supervised contrastive graph-summary pretraining method. Specifically, we generate and curate large-scale graph-summary pair data with the assistance of LLMs, and introduce a novel graph-summary pretraining method, combined with invariant learning, to enhance graph foundation models with strong cross-domain zero-shot transferability. For few-shot learning, we propose a novel graph prompt tuning technique aligned with our pretraining objective to mitigate catastrophic forgetting and minimize learning costs. Extensive experiments show the superiority of GraphCLIP in both zero-shot and few-shot settings, while evaluations across various downstream tasks confirm the versatility of GraphCLIP. Our code is available at: https://github.com/ZhuYun97/GraphCLIP.
Delivering superior search services is crucial for enhancing customer experience and driving revenue growth in e-commerce. Conventionally, search systems model user behaviors by combining user preference and query-item relevance statically, often through a fixed logical 'and' relationship. This paper reexamines existing approaches through a unified lens using causal graphs and Venn diagrams, uncovering two prevalent yet significant issues: entangled preference and relevance effects, and a collapsed modeling space. To surmount these challenges, our research introduces a novel framework, DRP, which enhances search accuracy through two components to reconstruct the behavior modeling space. Specifically, we implement preference editing to proactively remove the relevance effect from preference predictions, yielding untainted user preferences. Additionally, we employ adaptive fusion, which dynamically adjusts fusion criteria to align with the varying patterns of relevance and preference, facilitating more nuanced and tailored behavior predictions within the reconstructed modeling space. Empirical validation on two public datasets and a proprietary e-commerce search dataset underscores the superiority of our proposed methodology, demonstrating marked improvements in performance over existing approaches. The code is available at https://github.com/Applied-Machine-Learning-Lab/DRP.
SuiGPT MAD: Move AI Decompiler to Improve Transparency and Auditability on Non-Open-Source Blockchain Smart Contract
PDF ↗The vision of Web3 is to improve user control over data and assets, but one challenge that complicates this vision is the prevalence of non-transparent, scam-prone applications and vulnerable smart contracts that put Web3 users at risk. While code audits are one solution to this problem, the lack of smart contracts source code on many blockchain platforms, such as Sui, hinders the ease of auditing. A promising approach to this issue is the use of a decompiler to reverse-engineer smart contract bytecode. However, existing decompilers for Sui produce code that is difficult to understand and cannot be directly recompiled. To address this, we developed the SuiGPT Move AI Decompiler (MAD), a Large Language Model (LLM)-powered web application that decompiles smart contract bytecodes on Sui into logically correct, human-readable, and re-compilable source code with prompt engineering. Our evaluation shows that MAD's output successfully passes original unit tests and achieves a 73.33% recompilation success rate on real-world smart contracts. Additionally, newer models tend to deliver improved performance, suggesting that MAD's approach will become increasingly effective as LLMs continue to advance. In a user study involving 12 developers, we found that MAD significantly reduced the auditing workload compared to using traditional decompilers. Participants found MAD's outputs comparable to the original source code, improving accessibility for understanding and auditing non-open-source smart contracts. Through qualitative interviews with these developers and Web3 projects, we further discussed the strengths and concerns of MAD. MAD has practical implications for blockchain smart contract transparency, auditing, and education. It empowers users to easily and independently review and audit non-open-source smart contracts, fostering accountability and decentralization. Moreover, MAD's methodology could potentially extend to other smart contract languages, like Solidity, further enhancing Web3 transparency.
Frontier large language models (LLMs) have demonstrated remarkable performance across various knowledge-intensive enterprise tasks. However, these models are primarily trained on unstructured, general knowledge, which limits their effectiveness in domain-specific applications-particularly when tasks involve structured data sources or sensitive enterprise information. We propose the first Structured Knowledge for Large Language Models Workshop - SKnow-LLM, which aims to bridge this gap by promoting research on innovative methodologies and practical applications in this area. Through keynote talks, panel discussions and paper presentations, the workshop will foster in-depth discussions on recent advances, identify existing challenges, and explore promising directions for integrating structured knowledge into LLMs.
The demand for efficient Large Language Model (LLM) inference has surged with the rising adoption of Generative AI (GenAI) applications, particularly in areas such as agents and retrieval-augmented generation. Efficient inference serves two crucial purposes: it enables the deployment of LLM-centered applications that address critical business needs, while also facilitating rapid experimentation for researchers to extract valuable insights and new understandings. However, despite the field's rapid advancement and interdisciplinary nature, there remains a limited exchange of ideas and methodologies between production-facing practitioners and researchers seeking to experiment with new GenAI concepts quickly. To bridge this gap, we are introducing the first KDD workshop on Inference Optimization for Generative AI. Our goal is to create a collaborative platform where researchers and practitioners working across various use cases and stacks of efficient inference can come together to exchange research ideas, establish connections between different disciplines, and identify challenges and research questions that will shape future work.
The 3rd Workshop on Causal Inference and Machine Learning in Practice at KDD 2025 aims to bring together researchers, industry professionals, and practitioners to explore the application of causal inference within machine learning models. As causal machine learning techniques gain traction across industries, practical challenges related to trustworthiness, robustness, and fairness remain at the forefront. This workshop will provide a forum to discuss methodologies for evaluating causal models in real-world scenarios and explore innovative applications that integrate causal inference with generative AI (GenAI) and large language models (LLMs). Topics of interest include using GenAI and LLMs to facilitate causal inference tasks and leveraging causal inference techniques for evaluating and improving GenAI/LLM models. Building on the success of the previous workshop editions at KDD 2023 and KDD 2024, which attracted over 200 and 250 participants, respectively, this workshop will continue fostering collaboration between academia and industry. Through invited talks, contributed papers, and interactive discussions, we will address key challenges and opportunities at the intersection of causal inference and machine learning. As the field continues to evolve, this workshop serves as a crucial platform for knowledge exchange and innovation, driving forward the application of causal techniques in machine learning and AI.
The rapid deployment of Generative and Agentic AI systems-ranging from large language models to autonomous agents-has created a critical need for rigorous and trustworthy evaluation methodologies. As these models influence real-world decision-making, traditional performance metrics alone fall short in capturing issues of safety, ethical alignment, misinformation, and human-centered usability. This workshop addresses these challenges by fostering interdisciplinary discussions and innovations in evaluation strategies that go beyond conventional benchmarks. Topics include holistic and multi-perspective assessments, scalable evaluation pipelines, reasoning and goal alignment in agentic behavior, misinformation detection, cross-modal generation, and trust calibration. By advancing robust, user-centric, and societally grounded evaluation practices, this workshop contributes to expanding KDD's methodological frontier into the emerging domain of responsible AI systems.
The proposed ''SciSoc LLM Workshop: Large Language Models for Scientific and Societal Advances'' aims to explore the profound implications and potential of Large Language Models (LLMs) in driving forward scientific inquiry and addressing critical societal challenges. As LLMs such as GPT-4 continue to redefine boundaries in both complexity and capability, their integration into the scientific and societal domains is not just beneficial but essential. In particular, LLMs have demonstrated substantial value in improving our understanding of complex datasets and generating insights across various fields such as healthcare, environmental science, education, and public policy. By bringing together experts and enthusiasts from diverse fields, the workshop aims to foster a comprehensive understanding of how LLMs can redefine traditional research methodologies. Participants will explore innovative ways to harness the power of LLMs for greater efficiency and innovation in their respective fields, potentially catalyzing a new era of scientific and societal advancement.
Prompt engineering plays a critical role in enabling the effective use of large language models (LLMs). LLMs exhibit unpredictable sensitivity to various input factors that may result in a performance gap between prompts that are semantically indistinguishable. Therefore, LLM researchers and practitioners often optimize their prompts in an ad hoc manner due to the lack of systematic methods. Prompt optimization remains an open problem due to the rapidly evolving landscape of NLP tasks, target LLMs, and associated best practices. To address this gap, we propose the first KDD workshop on Prompt Optimization. This workshop aims to bring together researchers and practitioners working on prompt design, automatic optimization, and evaluation, fostering the exchange of ideas and methodologies. By exploring topics such as discrete and soft prompt tuning, low-resource applications, we aim to establish best practices, identify challenges, and drive future research in this critical area.
Model extraction attacks pose significant security threats to deployed language models, potentially compromising intellectual property and user privacy. This survey provides a comprehensive taxonomy of LLM-specific extraction attacks and defenses, categorizing attacks into functionality extraction, training data extraction, and prompt-targeted attacks. We analyze various attack methodologies including API-based knowledge distillation, direct querying, parameter recovery, and prompt stealing techniques that exploit transformer architectures. We then examine defense mechanisms organized into model protection, data privacy protection, and prompt-targeted strategies, evaluating their effectiveness across different deployment scenarios. We propose specialized metrics for evaluating both attack effectiveness and defense performance, addressing the specific challenges of generative language models. Through our analysis, we identify critical limitations in current approaches and propose promising research directions, including integrated attack methodologies and adaptive defense mechanisms that balance security with model utility. This work serves NLP researchers, ML engineers, and security professionals seeking to protect language models in production environments.
Recommender systems (RSs) have become essential for alleviating information overload and matching users with relevant content. Traditionally, RSs have focused on personalized content distribution, leveraging user interaction data and various features to rank and recommend existing items. Recently, diffusion models (DMs) have emerged as powerful generative paradigms, introducing new possibilities for RSs to not only enhance their performance for content distribution but also extend their capability boundaries to personalized content creation. On the one hand, DMs enhance the recommendation performance by mitigating challenges such as sparse user-item interactions, weak latent representations, and noisy data. On the other hand, DMs enable personalized content creation, transforming RSs from passive distributors into active generators of user-specific media assets, such as customized images, posters, and multimedia content. Given such a transformative paradigm shift, this survey provides a comprehensive review of the integration of diffusion models into recommender systems, exploring key methodologies, application scenarios, and their impact on recommendation effectiveness, diversity, and personalization. We categorize DM-based recommendation paradigms into content distribution and content creation, compare integration strategies, and discuss open challenges and future directions. By bridging the gap between diffusion models and recommender systems, this work aims to guide researchers and practitioners in developing the next generation of generative AI-powered recommendation solutions.
Spatio-Temporal (ST) data science, which includes sensing, managing, and mining large-scale data across space and time, is fundamental to understanding complex systems in domains such as urban computing, climate science, and intelligent transportation. Traditional deep learning approaches have significantly advanced this field, particularly in the stage of ST data mining. However, these models remain task-specific and often require extensive labeled data. Inspired by the success of Foundation Models (FM), especially large language models, researchers have begun exploring the concept of Spatio-Temporal Foundation Models (STFMs) to enhance adaptability and generalization across diverse ST tasks. Unlike prior architectures, STFMs empower the entire workflow of ST data science, ranging from data sensing, management, to mining, thereby offering a more holistic and scalable approach. Despite rapid progress, a systematic study of STFMs for ST data science remains lacking. This survey aims to provide a comprehensive review of STFMs, categorizing existing methodologies and identifying key research directions to advance ST general intelligence.
Synthetic data is often positioned as a solution to replace sensitive fixed-size datasets with a source of unlimited matching data, freed from privacy concerns. There has been much progress in synthetic data generation over the last decade, leveraging corresponding advances in machine learning and data analytics. In this survey, we cover the key developments and the main concepts in tabular synthetic data generation, including paradigms based on probabilistic graphical models and on deep learning. We provide background and motivation, before giving a technical deep-dive into the methodologies. We also address the limitations of synthetic data, by studying attacks that seek to retrieve information about the original sensitive data. Finally, we present extensions and open problems in this area.
AI shows promising potential to improve patient health outcomes, and the accelerating pace of technological advancement suggests health AI may be approaching a transformative threshold. Despite this promise, widespread implementation remains elusive. What barriers persist, and how might we chart a viable path forward? This talk examines two critical pathways to realizing AI's potential in healthcare: how can we enable large-scale access to clinical data and how can we evolve AI to earn the trust of medical professionals? The first challenge - lack of access to clinical data - is well-recognized but persistent. Despite recent increases in publicly available healthcare data, we still lack the volume and diversity needed to ensure accuracy across demographic groups and medical conditions. In other AI domains such as text generation, the breakthrough to reliable performance came through massive-scale training datasets. Healthcare requires a similar scale, yet patient privacy rightfully restricts data access. This presentation will explore methods for generating synthetic patient records that maintain privacy while providing the necessary training scale. The second pathway involves evolving AI to build trust within the clinical community. Medical AI must offer transparency in its decision-making processes. For instance, when answering whether ''Alice takes blood thinners'', an AI system must provide supporting evidence rather than a simple yes/no response. Two approaches will be presented that address this need: fact-verification systems for clinical claims in structured data and semantic highlighting for unstructured text. For predictive scenarios such as ''Will Alice need ventilator support in the next 48 hours?'', I will demonstrate how predictions coupled with counterfactual explanations enhance clinical trust. In addition, AI can earn provider trust by solving problems clinicians lack time to address, such as translating radiology reports into patient-friendly explanations. I will showcase methodologies that effectively bridge this communication gap. Clearing these pathways is essential to equip healthcare providers with AI tools that enable more accurate, efficient decision-making. Despite the challenges, there are compelling reasons for optimism that these barriers can be overcome, bringing us closer to truly AI-enabled healthcare.
IVMR suite: An Industrial-scale Virtual Machine Rescheduling Dataset and Benchmark for Elastic Cloud Service
PDF ↗Virtual Machine Rescheduling (VMR) plays a crucial role in maintaining service quality and resource efficiency in elastic cloud computing. However, existing datasets and benchmarks primarily focus on VM scheduling tasks, while lacking industrial-scale datasets and standardized evaluation for the more complex and crucial rescheduling problems. To address these challenges, we present IVMR suite, the first industrial-scale suite for VMR research comprising two core components: 1) IVMR-D, an industrial-grade VMR dataset mined from a real cloud data center, integrating complete resource specifications and complex operation constraints. The dataset is systematically structured based on data size and optimization objectives. 2) IVMR-B, a benchmark for the VMR problem that establishes seamless integration of consistent evaluation and the provision of baselines spanning optimization, metaheuristic, heuristic, and machine learning-based methodologies. Our comprehensive experimental evaluation demonstrates that all tested VMR algorithms struggle to effectively balance solution quality with computational efficiency while showing limited scalability across tasks with varying complexity levels. These findings emphasize the urgency of improving VMR algorithms for industrial deployments.
Capillary Dataset: A Dataset of Nail-fold Capillaries Captured by Microscopy for Diabetes Detection
PDF ↗Diabetes mellitus is a chronic condition marked by insufficient insulin utilization or production, causing metabolic dysregulation. If not controlled, it can cause serious complications that affect major organ systems such as the cardiovascular and ocular systems. Early diagnosis is essential for timely interventions to ensure effective glycemic control and lower the risk of complications. The present study introduced a robust and comprehensive dataset derived from nail-fold capillaroscopy. This dataset, which includes imaging and some video data from 126 individuals diagnosed with type 2 diabetes mellitus (T2DM) alongside 76 healthy controls, consisted of 3283 images obtained from diabetic subjects and 3412 images from non-diabetic participants. All images were acquired at 390× high magnification with a high resolution of 640×480 pixels, including corresponding video data, ensuring a thorough and detailed dataset for our study. The dataset was organized into two analytical tracks. The first track focused on nail-fold morphology, classifying images into four types: hairpin, crossing, tortuous, and bushy. After filtering out duplicates and low-quality images, 1279 images were selected for analysis. The second track involved binary classification for the detection of diabetes, differentiating healthy individuals from people with diabetes. This utilized the complete dataset and the augmented versions created by combining multiple images into composite formats to enhance feature representation. Additionally, we evaluated several benchmark deep-learning models such as Vision Transformers (ViT) and Convolutional Neural Networks (CNNs) for morphological classification and diabetes detection tasks. This analysis illuminated current model performance, highlighted challenges, and paved the way for future research opportunities. Importantly, this dataset is critical and holds significant potential for advancing non-invasive automated diagnostic methodologies in diabetes-related nail-fold capillary research, offering a promising future for diabetes research. The dataset is available at: https://huggingface.co/datasets/HanaNguyen/Capillary-Dataset. The github: https://github.com/urgonguyen/Capillarydataset.git.
Scientific researchers need intensive information about datasets to effectively evaluate and develop theories and methodologies. The information needs regarding datasets are implicitly embedded in particular research tasks, rather than explicitly expressed in search queries. However, existing scientific retrieval and question-answering (QA) datasets typically address straightforward questions, which do not align with the distribution of real-world research inquiries. To bridge this gap, we developed ScIRGen, a dataset generation framework for scientific QA & retrieval that more accurately reflects the information needs of professional science researchers, and uses it to create a large-scale scientific retrieval-augmented generation (RAG) dataset with realistic queries, datasets and papers. Technically, we designed a dataset-oriented information extraction method that leverages academic papers to augment the dataset representation. We then proposed a question generation framework by employing cognitive taxonomy to ensure the quality of synthesized questions. We also design a method to automatically filter synthetic answers based on the perplexity shift of LLMs, which is highly aligned with human judgment of answers' validity. Collectively, these methodologies culminated in the creation of the 61k QA dataset, ScIRGen-Geo. We benchmarked representative methods on the ScIRGen-Geo dataset for their question-answering and retrieval capabilities, finding out that current methods still suffer from reasoning from complex questions. This work advances the development of more sophisticated tools to support the intricate information needs of the scientific community.
Neurophysiologically Realistic Environment for Comparing Adaptive Deep Brain Stimulation Algorithms in Parkinson's Disease
PDF ↗Adaptive deep brain stimulation (aDBS) has emerged as a promising treatment for Parkinson's disease (PD). In aDBS, a surgically placed electrode sends dynamically altered stimuli to the brain based on neurophysiological feedback: an invasive gadget that limits the amount of data one could collect for optimizing the control offline. As a consequence, a plethora of synthetic models of PD and those of the control algorithms have been proposed. Herein, we introduce the first neurophysiologically realistic benchmark for comparing said models. Specifically, our methodology covers not only conventional basal ganglia circuit dynamics and pathological oscillations, but also captures 15 previously dismissed physiological attributes, such as signal instabilities and noise, neural drift, electrode conductance changes and individual variability - all modeled as spatially distributed and temporally registered features via beta-band activity in the brain and a feedback. Furthermore, we purposely built our framework as a structured environment for training and evaluating deep reinforcement learning (RL) algorithms, opening new possibilities for optimizing aDBS control strategies and inviting the machine learning community to contribute to the emerging field of intelligent neurostimulation interfaces. Code repository: https://github.com/NevVerVer/DBS-Gym
The success of clinical trials of longevity drugs relies heavily on identifying integrative health and aging biomarkers, such as biological age. Epigenetic aging clocks predict the biological age of an individual using their DNA methylation profiles, commonly retrieved from blood samples. However, there is no standardized methodology to validate and compare epigenetic clock models as yet. We propose ComputAgeBench, a unifying framework that comprises such a methodology and a dataset for comprehensive benchmarking of different clinically relevant aging clocks. Our methodology exploits the core idea that reliable aging clocks must be able to distinguish between healthy individuals and those with aging-accelerating conditions. Specifically, we collected and harmonized 66 public datasets of blood DNA methylation, covering 19 such conditions across different ages, and tested 13 published clock models. Additionally, we compiled 46 separate datasets to facilitate the training of new aging clocks. We believe our work will bring the fields of aging biology and machine learning closer together for the research on reliable biomarkers of health and aging. Code https://github.com/ComputationalAgingLab/ComputAge Dataset https://huggingface.co/datasets/computage/computage_bench