Multi-modal recommender systems (MMRS) have gained significant attention due to their ability to leverage information from various modalities to enhance recommendation quality. However, existing negative sampling techniques often struggle to effectively utilize the multi-modal data, leading to suboptimal performance. In this paper, we identify two key challenges in negative sampling for MMRS: (1) producing cohesive negative samples contrasting with positive samples and (2) maintaining a balanced influence across different modalities. To address these challenges, we propose NegGen, a novel framework that utilizes multi-modal large language models (MLLMs) to generate balanced and contrastive negative samples. We design three different prompt templates to enable NegGen to analyze and manipulate item attributes across multiple modalities, and then generate negative samples that introduce better supervision signals and ensure modality balance. Furthermore, NegGen employs a causal learning module to disentangle the effect of intervened key features and irrelevant item attributes, enabling fine-grained learning of user preferences. Extensive experiments on real-world datasets demonstrate the superior performance of NegGen compared to state-of-the-art methods in both negative sampling and multi-modal recommendation.
论文检索
输入标题、作者或关键词,从 1,620 篇学术成果中精准定位
Multimodal recommender systems enhance recommendation performance by integrating information from different modalities (e.g., text and images). A common approach is to link items with high modality similarity in modality graphs, helping users explore their interests more broadly. However, existing methods often introduce noise when enhancing modality graphs, making it challenging to effectively balance performance and accuracy. To address this issue, we propose an Interest Tree Augmented Modality Graph RecommendER for Multimodal Recommendation (TAMER). In this framework, we first redistribute item modality features using various component analysis methods to ensure more reliable item similarity within modality graphs. Next, we construct interest graphs based on reliable semantic relationships and prune the interest graphs into multiple interest trees. These interest trees are then applied to the multimodal item-item homogeneous graph to extend potential links within the modality homogeneous graph. The interest tree-based enhancement method effectively captures high-order relationships in the modality graph while avoiding noisy links. The effectiveness of the proposed method is demonstrated through comprehensive experiments on three real-world datasets. Compared with the strongest baseline methods, our method achieves an average improvement of 9.98% across four evaluation metrics. The source code is available at https://github.com/Z-last-ONE/TAMER.
We focus on the approximate nearest neighbor search (ANNS) in high dimensional space, which is a fundamental technique in computer vision and multimedia database. Among the ANNS solutions, graph-based approaches achieve excellent performance by executing a routing algorithm on a proximity graph to retrieve the nearest neighbors. However, most of their routing strategies are heuristic-based greedy routing, leading to suboptimal search results with large number of hops. In this paper, we propose a novel routing paradigm on graphs for ANNS problem by deep reinforcement learning. We design a reinforcement model to learn the routing policy by making use of both graph global and local topology information. A hops-optimized reward mechanism is devised to enable the model to be more efficient and effective. The final searching algorithm with the learned model is able to find the nearest neighbors without any backtracking in a small number of hops. Comprehensive experiments on real-world datasets demonstrate the superiorities of the proposed method over the state-of-the-art ANNS approaches.
Next Point-of-Interest (POI) recommendation aims to predict user's subsequent destinations based on historical check-in sequences, thereby enhancing travel experiences. While traditional methods primarily rely on unique identifiers (IDs) to represent POIs, they face data scarcity challenges. Recent multi-modal approaches offer alternatives but struggle with two key issues: inadequate handling of heterogeneity between ID and multi-modal features, and difficulties in unified framework integration, limiting their potential benefits. To address these limitations, we propose IM-POI, a novel framework that leverages the complementary strengths of both ID embeddings and multi-modal representations for next POI recommendation. In our framework, a global POI weighted transition graph inspired by TF-IDF captures sequential dependencies and enhances memorization capabilities, while a geographical graph incorporates spatial information into multi-modal features to be consistent with real-world visitation patterns. To address representation integration, we introduce an IM-Aligner module to prevent representation collapse during distribution matching. Extensive experiments on three real-world datasets demonstrate that IM-POI significantly outperforms state-of-the-art baselines.
Personalized product search (PPS) aims to retrieve products relevant to the given query considering user preferences within their purchase histories. Since large language models (LLM) exhibit impressive potential in content understanding and reasoning, current methods explore to leverage LLM to comprehend the complicated relationships among user, query and product to improve the search performance of PPS. Despite the progress, LLM-based PPS solutions merely take textual contents into consideration, neglecting multimodal contents which play a critical role for product search. Motivated by this, we propose a novel framework, HMPPS, for Harnessing Multimodal large language models (MLLM) to deal with Personalized Product Search based on multimodal contents. Nevertheless, the redundancy and noise in PPS input stand for a great challenge to apply MLLM for PPS, which not only misleads MLLM to generate inaccurate search results but also increases the computation expense of MLLM. To deal with this problem, we additionally design two query-aware refinement modules for HMPPS: 1) a perspective-guided summarization module that generates refined product descriptions around core perspectives relevant to search query, reducing noise and redundancy within textual contents; and 2) a two-stage training paradigm that introduces search query for user history filtering based on multimodal representations, capturing precise user preferences and decreasing the inference cost. Extensive experiments are conducted on four public datasets to demonstrate the effectiveness of HMPPS. Furthermore, HMPPS is deployed on an online search system with billion-level daily active users and achieves an evident gain in A/B testing.
Same-product identification serves as a critical infrastructure in e-commerce systems, enabling accurate product matching across heterogeneous marketing representations for key applications such as price comparison and personalized recommendation. Conventional approaches typically depend on manual feature engineering and extensive rule tuning, which limits their adaptability to varying identification criteria across different product categories and inconsistent business scenarios. To overcome these challenges, we propose an end-to-end same-product identification model powered by multimodal large language models (MLLMs) that inherently support multimodal alignment and exhibit strong generalization across diverse real-world settings. We first introduce a novel group-wise annotation pipeline to construct a high-quality dataset, consisting of diverse product pairs with multimodal presentations and labeled at the SKU level. Building on this dataset, we incorporate task-specific training recipes from the perspective of data augmentation, resulting in our SaP-Bot, which demonstrates advanced performance and generalization capabilities. Moreover, we identify a strong correlation between the output logits of MLLMs and the product similarity, enabling interpretable confidence estimation that benefits both data annotation and downstream applications.
In this paper, we present a novel Self-Supervised Learning (SSL) framework tailored for Multi-View Clustering (MVC), which learns cross-view semantic representations with clear clustering boundaries and derives balanced clustering in an end-to-end manner. Concretely, we propose a generative SSL module that learns high-level semantic representations by recovering randomly masked views from observed views. Then the extracted representations are unified via a sample-level local fusion mechanism and projected into a unit-hypersphere space with evenly distributed cluster prototypes such that the pseudo labels can be directly retrieved using cosine similarity. For each sample, we define highly credible positive pairs of the same cluster and negative pairs of different clusters and design a contrastive SSL module to force the sample to move toward its cluster prototype while farther from the other prototypes in the embedding space. Consequently, the representations exhibit clearer clustering boundaries, and the two SSL modules benefit each other. Finally, we further introduce a clustering regularizer to prevent trivial solutions and derive balanced clustering with theoretical guarantees. Comprehensive evaluations over eight benchmark datasets validate the effectiveness of our proposals against ten state-of-the-art MVC methods.
Unsupervised hashing is applied in large-scale multimodal retrieval by mapping original data from heterogeneous modalities into compact binary codes. Transformer-based retrieval augmented generation possesses significant advantages in retrieval accuracy and context-awareness, yet faces scalability challenges due to the computational overhead of dense embedding. Thus, the integration of hash learning and Transformer provides a feasible improvement scheme, which can achieve efficient retrieval preserving semantic association. This paper proposes a novel Unsupervised Similarity-Fusion Transformer Hashing for multimodal retrieval, denoted as USFTH. Initially, the modal fusion similarity matrix based on Gaussian kernel, sigmoid function, and Laplacian transformation is introduced to construct a discriminative similarity matrix, ensuring that semantic correlation among samples can be captured precisely. Then, cross-modal multiplex joint construction via Transformer-based attention mechanisms is designed, realizing effective integration of heterogeneous modalities in the similarity matrix through multi-path fusion. Furthermore, the consensus fusion strategy is proposed to ensure that hash codes generated under unsupervised conditions possess a uniform distribution and achieve accurate retrieval. In addition, comprehensive experiments on MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets demonstrate the superior performance of USFTH to state-of-the-art hashing approaches.
Recently, prompt learning has achieved remarkable success in adapting pre-trained Vision-Language Models (VLMs) to downstream tasks such as image classification. However, its application to the downstream Image-Text Retrieval (ITR) task is more challenging. We find that the challenge lies in discriminating both fine-grained attributes and similar subcategories of the downstream data. To address this challenge, we propose Dual prompt Learning with Joint Category-Attribute Reweighting (DCAR), a novel dual-prompt learning framework to achieve precise image-text matching. The framework dynamically adjusts prompt vectors from both semantic and visual dimensions to improve the performance of CLIP on the downstream ITR task. Based on the prompt paradigm, DCAR jointly optimizes attribute and category features to enhance fine-grained representation learning. Specifically, (1) at the attribute level, it dynamically updates the weights of attribute descriptions based on text-image mutual information correlation; and (2) at the category level, it introduces negative samples from multiple perspectives with category-matching weighting to learn subcategory distinctions. To validate our method, we construct the Fine-class Described Retrieval Dataset (FDRD), which serves as a challenging benchmark for ITR in downstream data domains. It covers over 1,500 downstream fine categories and 230,000 image-caption pairs with detailed attribute annotations. Extensive experiments on FDRD demonstrate that DCAR achieves state-of-the-art performance over existing baselines. The code and data are available at https://github.com/wyf202322/DCAR.
Crowd Dynamics Demand Adaptivity: Self-Adaptive Physics-Informed Neural Network for Crowd Simulation
Crowd simulation is crucial for urban planning, traffic management, public safety, and immersive environments. A fundamental challenge is capturing adaptive human behaviors that evolve dynamically with social interactions and task demands. Recently, physics-informed neural networks (PINNs) seamlessly integrate interpretable physics-based models with flexible data-driven learning, significantly enhancing simulation realism. However, current PINN-based methods typically rely on rigid representations of pedestrian perceptions and static task priorities of motion planning, limiting their ability to capture real-world social complexities and behavioral adaptability. To this end, we introduce SA-PINN, a novel Self-Adaptive Physics-Informed Neural Network specifically designed for modeling adaptive crowd behaviors. SA-PINN features two innovative adaptive modules: a self-adaptive social perception module, guided by a visual-field physics model to capture context-dependent social interactions dynamically; and a self-adaptive multi-task PINN training module, automatically balancing key motion objectives such as goal-reaching, collision avoidance, and alignment with real data. By jointly enabling perception-level and task-level adaptations within a unified physics-informed framework, SA-PINN generates highly realistic and physically consistent crowd simulations across diverse environmental contexts. Comprehensive evaluations on three real-world datasets (Lane, Cross 90, and GC) reveal that SA-PINN achieves a 29.7% gain in microscopic trajectory accuracy and enhances macroscopic density similarity by 23.5% compared to the best-performing baselines.
Public response prediction is critical for understanding how individuals or groups might react to specific events, policies, or social phenomena, making it highly valuable for crisis management, policy-making, and social media analysis. However, existing works face notable limitations. First, they lack micro-level personalization, producing generic responses that ignore individual user preferences. Moreover, they overlook macro-level sentiment distribution and only deal with individual-level sentiment, constraining them from analyzing broader societal trends and group sentiment dynamics. To address these challenges, we propose SocialAlign, a unified framework that predicts real-world responses at both micro and macro levels in social contexts. At the micro level, SocialAlign employs SocialLLM with an articulate Personalized Analyze-Compose LoRA (PAC-LoRA) structure, which deploys specialized expert modules for content analysis and response generation across diverse topics and user profiles, enabling the generation of personalized comments with corresponding sentiments. At the macro level, it models group sentiment distributions and aligns predictions with real-world sentiment trends derived from social media data. To evaluate SocialAlign in real-world scenarios, we introduce SentiWeibo, a large-scale dataset curated from authentic social interactions on the Weibo platform. Experimental results on our SentiWeibo and related LaMP benchmark demonstrate that SocialAlign surpasses strong baselines, showing improved accuracy, interpretability, and generalization in public response prediction. We hope our work inspires further research in public response prediction and computational social science: https://github.com/Znull-1220/SocialAlign.
Multimodal Sentiment Analysis (MSA) aims to identify sentiment polarity and intensity in media. Current methods typically employ a two-stage pipeline: extracting features from each modality, then predicting sentiment based on fused representations. However, most fusion strategies align features from different modalities in a single step, leading to conflicts during cross-modal interactions and hindering the modeling of hierarchical sentiment dependencies. Additionally, existing methods often overlook the dominant role of textual modality in high level latent fusion space, causing explicit linguistic sentiment cues to be obscured by redundant information. To address these issues, DDSE (Decoupled Dual-Stream Enhanced framework) is proposed in this work, which decouples features into public and private representations for improved feature enhancement and cross-modal interaction. The proposed TC-Mamba module enables progressive cross-modal interactions within shared state transition matrices under a text-guided fusion paradigm, effectively preserving sentiment cues and minimizing redundancy. Additionally, DDSE adopts a multi-task learning strategy to further enhance overall performance. Extensive experiments on the MOSI and MOSEI datasets demonstrate that DDSE achieves state-of-the-art results, with Acc-5 improvements of 3.06% and 0.1%, respectively, underscoring its effectiveness in MSA. Ablation studies confirm the critical contributions of each component within the framework. Code is available at https://anonymous.4open.science/r/DDSE-76D6.
This paper explores silent speech decoding in active brain-computer interface (BCI) systems, which offer more natural and flexible communication than traditional BCI applications. We collected a new silent speech dataset of over 120 hours of electroencephalogram (EEG) recordings from 12 subjects, capturing 24 commonly used English words for language model pretraining and decoding. Following the recent success of pretraining large models with self-supervised paradigms to enhance EEG classification performance, we propose Large Brain Language Model (LBLM) pretrained to decode silent speech for active BCI. To pretrain LBLM, we propose Future Spectro-Temporal Prediction (FSTP) pretraining paradigm to learn effective representations from unlabeled EEG data. Unlike existing EEG pretraining methods that mainly follow a masked-reconstruction paradigm, our proposed FSTP method employs autoregressive modeling in temporal and frequency domains to capture both temporal and spectral dependencies from EEG signals. After pretraining, we finetune our LBLM on downstream tasks, including word-level and semantic-level classification. Extensive experiments demonstrate significant performance gains of the LBLM over fully-supervised and pretrained baseline models. For instance, in the difficult cross-session setting, our model achieves 47.2% accuracy on semantic-level classification and 42.3% in word-level classification, outperforming baseline methods substantially. Our research advances silent speech decoding in active BCI systems, offering an innovative solution for EEG language model pretraining and a new dataset for fundamental research.
Learning discriminative micro-expression (ME) features from low-intensity facial movements is a key challenge for micro-expression recognition (MER). Although existing research has demonstrated that the appearance, motion and geometric information are distinguishing for MEs, the effectiveness of merging these information is still unclear. Thus, this paper proposes a Multi-information Hierarchical Fusion Transformer (MiHF-Tr) model to fully and effectively aggregate the facial appearance, motion, and geometric information of MEs, exploring a more reasonable way of multi-information fusion. As different information is homology, MiHF-Tr introduces a local and global hierarchy fusion framework to fuse them by modeling their local and global semantic consistency. Considering the bias of different information in feature representation ability, a single-core self-attention is proposed to achieve local multi-information fusion, which focuses on strong information and supplements it with weak information. The experimental results demonstrate that the fusion of appearance, motion, and geometric features is discriminative, and the proposed method can effectively aggregate multiple information, achieving competitive performance.
Micro-expression analysis (MEA) is crucial for detecting subtle emotional cues, with applications in lie detection and psychological assessment. Existing methods struggle with three main challenges: 1) Noise sensitivity arising from the inherent subtlety of micro-expressions. 2) Reliance on fixed priors and apex annotations. 3) Information redundancy, with static features often dominating over dynamic emotional cues. To address these challenges, we propose Ac4AU, a framework inspired by Regulatory Focus Theory (RFT) that utilizes structured representation learning to decompose dynamic emotional patterns from redundant features. Specifically, AC4AU first leverages a face recognition backbone to extract robust yet redundant static representations. Secondly, a Frequency-aware Redundancy Decomposer (FRD) is introduced to eliminate the Direct Current component and retain the dynamic and process-sensitive features. Finally, a dynamic expert allocation mechanism, embodied by the AU-specific Expert Router (AUsER), is adopted to learn localized facial motion patterns and capture long-term relationships, enabling AU-targeted supervision and enhancing generalization across diverse datasets. Rigorous experiments demonstrate that the apex-free AC4AU achieves performance comparable to state-of-the-art apex-dependent methods. Additionally, we conduct a statistical analysis that provides insights into the AU dependencies. Code will be made available upon request.
The detection of telecom fraud faces significant challenges due to the lack of high-quality multimodal training data that integrates audio signals with reasoning-oriented textual analysis. To address this gap, we present TeleAntiFraud-28k, the first open-source audio-text slow-thinking dataset specifically designed for automated telecom fraud analysis. Our dataset is constructed through three strategies: (1) Privacy-preserved text-truth sample generation using automatically speech recognition-transcribed call recordings (with anonymized original audio), ensuring real-world consistency through text-to-speech model regeneration; (2) Semantic enhancement via large language model based self-instruction sampling on authentic ASR outputs to expand scenario coverage; (3) Multi-agent adversarial synthesis, which simulates emerging fraud tactics through predefined communication scenarios and fraud typologies, enriches the conversation samples. The generated dataset contains 28,511 rigorously processed audio-text pairs with a total audio duration of more than 307 hours, complete with detailed annotations for fraud reasoning. The dataset is divided into three tasks: scenario classification, fraud detection, fraud type classification. Furthermore, we construct TeleAntiFraud-Bench, a standardized evaluation benchmark comprising proportionally sampled instances from TeleAntiFraud-28k, to facilitate systematic testing of model performance, reasoning capabilities, and thought processes on telecom fraud detection tasks. We also contribute a supervised fine-tuning model based on Qwen2-Audio, trained on the TeleAntiFraud-28k training set, while open-sourcing the data processing framework to enable community-driven dataset expansion. This work establishes a foundational framework for multimodal anti-fraud research while addressing critical challenges in data privacy and scenario diversity. The code of this paper is publicly available at https://github.com/JimmyMa99/TeleAntiFraud.
Emotion recognition, as a core technology in mental health monitoring, has long been constrained by the intrusive nature of data collection methods relying on physiological signals and behavioral cues. Although existing motion-based approaches enable non-intrusive data acquisition, they often overlook the societal dimensions inherent in human behavior. As a result, they often exhibit a significant performance drop in real-world scenarios compared to laboratory settings. In this study, we analyzed the spatial distribution of participants' spatiotemporal trajectories and their visited Points of Interest (POIs), and observed significant differences under varying emotional states. Building on this observation, we propose a novel emotion recognition framework, SE2E, which innovatively incorporates the semantic information of POIs into the emotion recognition task. Specifically, SE2E employs a category-aware semantic embedding mechanism combined with a masked prediction task to ensure that the POI embeddings capture both categorical semantics and contextual information. It then structurally represents individual societal event patterns through a personalized spatiotemporal flow. Finally, a temporal-region consistency attention module is employed to extract continuous representations of societal events, thereby enabling a robust mapping from societal behavior to emotional state. Extensive experimental results demonstrate that SE2E outperforms state-of-the-art methods across multiple benchmarks. To the best of our knowledge, this is the first study to leverage societal event for emotion recognition, offering a new technical direction, benchmark, and insight for future research in the field.
Emotion recognition using electroencephalography (EEG) signals has attracted increasing attention in recent years. However, existing methods often lack generalization in cross-corpus settings, where a model trained on one dataset is directly applied to another without retraining, due to differences in data distribution and recording conditions. To tackle the challenge of cross-corpus EEG-based emotion recognition, we propose a novel framework termed Soft Contrastive Masked Modeling (SCMM). Grounded in the theory of emotional continuity, SCMM integrates soft contrastive learning with a hybrid masking strategy to effectively capture emotion dynamics (refer to short-term continuity). Specifically, in the self-supervised learning stage, we propose a soft weighting mechanism that assigns similarity scores to sample pairs, enabling fine-grained modeling of emotional transitions and capturing the temporal continuity of human emotions. To further enhance representation learning, we design a similarity-aware aggregator that fuses complementary information from semantically related samples based on pairwise similarities, thereby improving feature expressiveness and reconstruction quality. This dual design contributes to a more discriminative and transferable representation, which is crucial for robust cross-corpus generalization. Extensive experiments on the SEED, SEED-IV, and DEAP datasets show that SCMM achieves state-of-the-art (SOTA) performance, outperforming the second-best method by an average accuracy of 4.26% under both same-class and different-class cross-corpus settings. The source code is available at https://github.com/Kyler-RL/SCMM.
Large Language Model (LLM) agents have demonstrated impressive capabilities in social deduction games (SDGs) like Werewolf, where strategic reasoning and social deception are essential. However, current approaches remain limited to textual information, ignoring crucial multimodal cues such as facial expressions and tone of voice that humans naturally use to communicate. Moreover, existing SDG agents primarily focus on inferring other players' identities without modeling how others perceive themselves or fellow players. To address these limitations, we use One Night Ultimate Werewolf (ONUW) as a testbed and present MultiMind, the first framework integrating multimodal information into SDG agents. MultiMind processes facial expressions and vocal tones alongside verbal content, while employing a Theory of Mind (ToM) model to represent each player's suspicion levels toward others. By combining this ToM model with Monte Carlo Tree Search (MCTS), our agent identifies communication strategies that minimize suspicion directed at itself. Through comprehensive evaluation in both agent-versus-agent simulations and studies with human players, we demonstrate MultiMind's superior performance in gameplay. Our work presents a significant advancement toward LLM agents capable of human-like social reasoning across multimodal domains. Our code is available at https://github.com/CjangCjengh/onuw.
Multimodal large language models (MLLMs) have been widely applied across various fields due to their powerful perceptual and reasoning capabilities. In the realm of psychology, these models hold promise for a deeper understanding of human emotions and behaviors. However, recent research primarily focuses on enhancing their emotion recognition abilities, leaving the substantial potential in emotion reasoning, which is crucial for improving the naturalness and effectiveness of human-machine interactions. Therefore, in this paper, we introduce a multi-turn multimodal emotion understanding and reasoning (MTMEUR) benchmark, which encompasses 1,451 video data from real-life scenarios, along with 5,101 progressive questions. These questions cover various aspects, including emotion recognition, potential causes of emotions, future action prediction, etc. Besides, we propose a multi-agent framework, where each agent specializes in a specific aspect, such as background context, character dynamics, and event details, to improve the system's reasoning capabilities. Furthermore, we conduct experiments with existing MLLMs and our agent-based method on the proposed benchmark, revealing that most models face significant challenges with this task.