Multimodal Empathetic Response Generation (MERG) is crucial for building emotionally intelligent human-computer interactions. Although large language models (LLMs) have improved text-based ERG, challenges remain in handling multimodal emotional content and maintaining identity consistency. Thus, we propose E3RG, an Explicit Emotion-driven Empathetic Response Generation System based on multimodal LLMs which decomposes MERG task into three parts: multimodal empathy understanding, empathy memory retrieval, and multimodal response generation. By integrating advanced expressive speech and video generative models, E3RG delivers natural, emotionally rich, and identity-consistent responses without extra training. Experiments validate the superiority of our system on both zero-shot and few-shot settings, securing Top-1 position in the Avatar-based Multimodal Empathy Challenge on ACM MM'25. Our code is available at https://github.com/RH-Lin/E3RG.
论文检索
输入标题、作者或关键词,从 208 篇学术成果中精准定位
As embodied intelligence has become a new hot topic in current artificial intelligence research, facial reaction generation has increasingly become a key technology for achieving natural human-computer interaction. Existing methods typically rely on bimodal inputs of audio and visual signals, but they still suffer from poor cross-modal consistency and insufficient feature fusion, making it difficult to effectively capture complex facial features. To enhance the representational capacity of fused features, this paper proposes a Scattering-Conditioned Diffusion Model (SC-Diff), which extracts stable multi-scale structural features in the frequency domain via wavelet scattering transform and injects them into the diffusion generation process as conditional information, thereby enhancing the modeling ability of representing local facial variations. Furthermore, considering that different prediction tasks exhibit varying sensitivity to target changes during training, we introduce an uncertainty-based adaptive loss weighting strategy to dynamically balance three types of supervision targets: facial action units, facial affect, and facial expressions. Experimental results on the REACT 2025 dataset demonstrate that the proposed method outperforms existing state-of-the-art approaches across multiple evaluation metrics.
Accurately identifying personality traits is of profound significance for gaining in-depth insights into human behavior, facilitating efficient human-computer interaction, and developing personalized intelligent systems. However, existing studies often treat personality trait prediction and emotion recognition as relatively independent tasks, neglecting the inherent correlation between them. This report proposes a multimodal fusion prediction framework through the research topic ''On the Interaction between Personality and Emotion in Human Behavior and Social Interaction''. The core goal of this framework is to explore and verify the positive gain of emotional analysis on the accuracy of personality prediction. We extracted visual and audio features based on large-scale data pre-training and fine-tuning, aggregated video features at the character level to enhance the stability of personality prediction, and then integrated the video-level emotion prediction branch. By jointly optimizing the losses of personality prediction and emotion prediction, the generalization performance of the model is improved. Experimental results show that the multi-task learning method integrating emotional information can improve the prediction performance of personality traits to a certain extent, and achieved the first place in the MER-PR validation set of the MER2025 Challenge. This provides empirical evidence for us to explore the complex interaction between emotion and personality.
The Human-Object Interaction (HOI) task explores the dynamic interactions between humans and objects in physical environments, providing essential biomechanical and cognitive-behavioral foundations for fields such as robotics, virtual reality, and human-computer interaction. However, existing HOI data sets focus on details of affordance, often neglecting the influence of physical properties of objects on human long-term motion. To bridge this gap, we introduce the PA-HOI Motion Capture dataset, which highlights the impact of objects' physical attributes on human motion dynamics, including human posture, moving velocity, and other motion characteristics. The dataset comprises 562 motion sequences of human-object interactions, with each sequence performed by subjects of different genders interacting with 35 3D objects that vary in size, shape, and weight. This dataset stands out by significantly extending the scope of existing ones for understanding how the physical attributes of different objects influence human posture, speed, motion scale, and interacting strategies. We further demonstrate the applicability of the PA-HOI dataset by integrating it with existing motion generation methods, validating its capacity to transfer realistic physical awareness.
Accurate and robust 3D hand pose estimation (HPE) plays a crucial role in human-computer interaction. Existing 3D HPE solutions predominantly rely on vision-based or inertial measurement units (IMUs)-based methods. Vision-based methods benefit from rich appearance information for high-accuracy HPE but are sensitive to field of view (FoV), occlusion, motion blur and lighting. IMU-based methods can operate immune to optical sensitivity and FoV constraints but remain vulnerable to cumulative integration errors and drift. Given their complementary strengths, combining dual modalities offers a promising direction for HPE in complex environments. However, the lack of large-scale visual-inertial datasets has limited progress in this area. In this paper, we construct VIHand, the first large-scale glove-worn dataset for visual-inertial HPE, comprising over 1.4 million synchronized RGB-D and IMU frames from 15 subjects. It enables comprehensive research in HPE tasks, such as multimodal fusion and cross-modal knowledge transfer. Building on VIHand, we propose visual-inertial fusion network (VIFNet) for dual-modalities estimation, and its distilled student model (VIFNet-S) for IMU-only evaluation. Experimental results reveal that integrating visual and inertial modalities significantly improves the accuracy and robustness of 3D HPE, particularly under occlusion and motion blur. In IMU-only inference even sparse IMU configurations, models distilled from visual-inertial supervision achieve substantial performance gains, enabling robust HPE for challenging optical sensitive scenarios. Our dataset and supplementary materials are available on the project website: https://shirley0118.github.io/VIHand.
Nowadays, haptic data has gained a fast-growing volume with enormous interaction points during human-computer interaction and embodied AI. In the near future, the massive haptic signals -encompassing both kinesthetic and vibrotactile signals- will place significant demands on both communication and computing resources. To address this challenge, we propose the first task-oriented sematic codec of low-delay vibrotactile transmission, namely, vibrotactile semantic codec (VTSC). Specifically, we design a perception-based vibrotactile semantic extraction mechanism (PSEM) that considers the high and low thresholds of vibrotactile perception in effective semantic coding while adhering to the low delay constraint. Inspired by this principle, we then propose a vibrotactile semantic encoder (VSE) with local and global semantic extractors, which can efficiently extract and preserve semantic features within the short frame context. Besides, we present a semantic distribution loss function to enhance the learning of meaningful representations. Comprehensive experiments demonstrate the superiority of our VTSC, achieving significantly higher task accuracy than the state-of-the-art vibrotactile codecs at the same compression ratio (CR), e.g. 60% improvement when CR=256. When compared to transferred audio-visual sematic codecs, our VTSC also shows promising improvements, validating the effectiveness our approach.
High-quality thermal facial data is essential for advancing biometric recognition, surveillance, in-cabin driver monitoring, and human-computer interaction, all of which are integral for modern multimedia and interactive AI systems. In this work, we optimized the FLUX text-to-image diffusion model on diverse real-world thermal facial datasets to generate hyper-realistic 2D thermal facial images for both males and females, and propose a new dataset, ThermVision. To enhance their multimedia applicability, these images are processed through a video retargeting pipeline, where driving videos animate realistic facial expressions and head pose variations from a single 2D thermal image, producing high-fidelity thermal facial video sequences. The overall rendered dataset incorporates smart transformations, ensuring diversity across gender balance, extreme head pose variations, expressive facial dynamics, and facial accessories, making it a valuable resource for real-world applications. Additionally, we provide facial detection annotations to facilitate precise feature extraction and thermal-face analysis. To validate our synthetic dataset, we evaluate its effectiveness in thermal gender classification, as downstream machine learning task, along with thermal face localization and facial landmarks detection demonstrating its applicability in real-world scenarios. This approach significantly improves the availability, realism, and integration of thermal facial data, paving the way for more robust and immersive AI-powered thermal imaging applications. The dataset, code and associated models are available at- https://mali-farooq.github.io/ThermVision/
Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the rhythmic or semantic triggers from audio for generating contextualized gesture patterns and achieving pixel-level realism. To address these challenges, we introduce Contextual Gesture, a framework that improves co-speech gesture video generation through three innovative components: (1) a chronological speech-gesture alignment that temporally connects two modalities, (2) a contextualized gesture tokenization that incorporate speech context into motion pattern representation through distillation, and (3) a structure-aware refinement module that employs edge connection to link gesture keypoints to improve video generation. Our extensive experiments demonstrate that Contextual Gesture not only produces realistic and speech-aligned gesture videos but also supports long-sequence generation and video gesture editing applications, shown in Fig.1
In computer animation, game design, and human-computer interaction, synthesizing human motion that aligns with user intent remains a significant challenge. Existing methods have notable limitations: textual approaches offer high-level semantic guidance but struggle to describe complex actions accurately; trajectory-based techniques provide intuitive global motion direction yet often fall short in generating precise or customized character movements; and anchor poses-guided methods are typically confined to synthesize only simple motion patterns. To generate more controllable and precise human motions, we propose ProMoGen (Progressive Motion Generation), a novel framework that integrates trajectory guidance with sparse anchor motion control. Global trajectories ensure consistency in spatial direction and displacement, while sparse anchor motions only deliver precise action guidance without displacement. This decoupling enables independent refinement of both aspects, resulting in a more controllable, high-fidelity, and sophisticated motion synthesis. ProMoGen supports both dual and single control paradigms within a unified training process. Moreover, we recognize that direct learning from sparse motions is inherently unstable, we introduce SAP-CL (Sparse Anchor Posture Curriculum Learning), a curriculum learning strategy that progressively adjusts the number of anchors used for guidance, thereby enabling more precise and stable convergence. Extensive experiments demonstrate that ProMoGen excels in synthesizing vivid and diverse motions guided by predefined trajectory and arbitrary anchor frames. Our approach seamlessly integrates personalized motion with structured guidance, significantly outperforming state-of-the-art methods across multiple control scenarios.
Human activity recognition (HAR) is an evolving technique that offers innovative solutions across various domains, such as healthcare, sports training, and human-computer interactions. This paper addresses the novel challenge of video-based activity recognition, focusing on detecting and classifying athletes' actions to enable precision sports training. Conventional HAR methods based on direct video analysis incur excessive computational overhead and constrained applicability. In contrast, our novel transformer-based framework, namely RSFomer, converts videos into multivariate time series, and then detects and classifies the athletes' actions. However, sports videos often suffer from severe occlusion, which introduces significant noise to the converted time series and thus deteriorates recognition performance. To address this challenge, we implement several innovative strategies to improve the robustness of our framework. First, we propose a dual-scale filtering mechanism that leverages the unscented Kalman filter and kinematic constraints to reduce noise and outliers in the converted time series. Second, we incorporate the masking mechanism and temporal slicing mechanism to enhance the transformer's ability to handle anomalies and extract multi-scale features for accurate action recognition. We perform extensive evaluations on our Boxing dataset as well as the UEA and FineGym datasets. The results demonstrate that our RSFomer is effective, outperforming existing state-of-the-art methods with significant advantages.
The detection of hand contact states, which involves identifying interactions between hands and objects or other entities, is essential for the development of human-computer interaction systems and the comprehension of social dynamics. Previous approaches have made progress in modeling hand-object interactions. Nonetheless, they neglect critical cues between their hands and bodies, as well as those of others, thus constraining their ability to accurately detect interpersonal contact. The task remains challenging due to frequent occlusions, especially in crowded multi-person scenarios with complex contexts. In this paper, a novel hand-object-person interaction network, called HOPNet, is proposed to model contextual information between hands and objects, as well as between hands and bodies. Specifically, HOPNet consists of two components: (i) the Hand-Object Relation (HOR) module analyzes interaction patterns between hands and objects, capturing spatial and semantic relationships; (ii) the Contrastive Spatial Refinement (CSR) module learns hand-body interactions through contrastive geometric embedding and relative spatial enhancement, improving interpersonal contact recognition in crowded scenarios. Experiments on ContactHands and 100DOH datasets demonstrate that HOPNet outperforms state-of-the-art methods.
3D hand pose estimation has garnered great attention in recent years due to its critical applications in human-computer interaction, virtual reality, and related fields. Accurate estimation of hand joints is essential for high-quality hand pose estimation. However, existing methods neglect the importance of Distal Phalanx Tip (TIP) and Wrist in predicting hand joints overall and often fail to account for the phenomenon of error accumulation for distal joints in gesture estimation, which can cause certain joints to incur larger errors, resulting in misalignments and artifacts in pose estimation and degrading the overall reconstruction quality. To address this challenge, we propose a novel segmented architecture for enhanced hand pose estimation (EHPE). We perform a local extraction of the TIP and wrist, thus alleviating the effect of error accumulation on the prediction of the TIP and further reduce the predictive errors for all joints on this basis. EHPE consists of two key stages: In the TIP and Wrist Joints Extraction stage (TW-stage), the positions of the TIP and wrist joints are estimated to provide an initial accurate joint configuration; In the Prior Guided Joints Estimation stage (PG-stage), a dual-branch interaction network is employed to refine the positions of the remaining joints. Extensive experiments on two widely used benchmarks demonstrate that EHPE achieves state-of-the-art performance.
Emotion recognition based on electroencephalogram (EEG) aims to recognize emotional states for improving user experience in Human-Computer Interaction, often using subjects' responses as labels. Physiological signals on widely used datasets (i.e. SEED, SEED-IV, and DEAP) are collected during subjects watching different types of movies as video stimulus. As a result, when subjects' emotional states change from one to another during a single video stimulus, two challenges are inevitable for reliable emotion recognition due to dynamic emotional fluctuations: (1) inaccurate annotation of EEG data; (2) feature confusion in classifier boundaries from similar emotional states (e.g., low-intensity happiness and neutral). However, previous studies have not given sufficient attention to the impact of the dynamic emotional fluctuations, leading to unreliable emotion recognition, especially in cross-subject emotion recognition scenarios. In this paper, we propose a Prototypes Collaborative Learning with Consistency Awareness (PCLCA) method to improve the reliability of cross-subject emotion recognition by introducing prototype learning. Specifically, a consistency awareness mechanism is designed to compute the consistency between labels and actual emotional states. Furthermore, a prototype collaborative strategy is adopted to adaptively estimate the uncertainty of model predictions by computing the similarity between features and prototypes. Extensive experiments on three benchmark datasets demonstrate that PCLCA effectively alleviates label noise and reduces uncertain predictions, outperforming existing baseline models.
Real-time emotion recognition provides promising applications for mental healthcare monitoring and human-computer interaction design. Electroencephalography (EEG) emotion recognition has become a hot topic in the field of affective computing and intelligent brain-computer interface (BCI), and it is a feasible solution for achieving real-time emotion recognition. However, due to the uncertainty and individual specificity of emotional cognition, there are still some challenges in achieving efficient online emotion decoding applications. To address this, in this work, we propose an online emotion decoding method named DMSGL (Real-Time EEG Emotion Recognition from Dynamic Mixed Spatiotemporal Graph Learning). Specifically, in the DMSGL, we propose to explore the latent emotion-related graph features from EEG with cognition-inspired and data-driven learning strategies, and the temporal analysis with attention learning is utilized to further extract the robust spatiotemporal graph patterns for efficient EEG emotion decoding. Both simulated online emotion decoding and real-time emotion monitoring experimental results have consistently indicated that the proposed DMSGL can effectively satisfy the application requirements of real-time emotion decoding and achieves an accuracy of 68.35% in real-world online scenarios. Compared with other baseline methods, the proposed DMSGL has improved by 2-5% in the scenario of real-time emotion recognition. In conclusion, the proposed DMSGL provides a promising solution for realizing real-time emotion recognition and further exploring related applications. Our code is released on https://github.com/UESTC-BAC/DMSGL.
Emotion recognition based on multimodal physiological signals is playing an increasingly important role in areas such as human-computer interaction and disease diagnosis, attracting growing attention from the research community. Current studies primarily focus on emotion recognition under unified data collection paradigms, overlooking the prevalent issue of imperfect modality matching in real-world scenarios. In particular, existing methods fail to effectively utilize these mismatched modalities, leading to incomplete emotional representations. This limits the model's ability to accurately capture the multidimensional semantic features of emotions, thereby constraining its effectiveness and applicability in practical settings. To address this challenge, we propose MoCERNet. At the modality level, it first reduces the domain gap among matched modalities and then aligns mismatched modalities in a semantics-aware manner, guided by the matched ones. At the decision level, it further mitigates the global distribution discrepancies to achieve a more complete emotional representation. In addition, we design a Nervous System Functional Structure Transformer (NFSformer) that enables the model to focus on the correlation between different brain regions and peripheral physiological signals under various emotional states, thereby enhancing its capacity to model complex emotional processes. Experiments on three multimodal emotion datasets demonstrate that MoCERNet outperforms state-of-the-art baselines under imperfect modality matching scenarios.
Micro-expressions (MEs) are involuntary facial expressions that reveal genuine emotions and have significant applications in fields such as psychology, security, and human-computer interaction. However, previous ME datasets are mainly collected in controlled laboratory environments, such as fixed views, single illumination and head movements, limited subjects and the lack of background. There are significant gaps between them and the real world. To handle this issue, we introduce a novel Natural Micro-Expression (NaME) dataset, a natural dataset collected under unconstrained real-world conditions. It encompasses (1) diverse subjects, multiple views and varying head movements ; (2) rich background information, providing a more realistic benchmark for the micro-expression recognition (MER) research. Furthermore, we propose a MER benchmark for natural environments, named MixFormer. MixFormer includes an efficient sparse attention mechanism to capture subtle facial motions from various factors, and a face-background mix of attention module to model the environment context to help MER. Extensive experiments are conducted to analyze our NaME dataset and benchmark. We believe that our dataset and benchmark will pave the way for future research in MER beyond controlled settings, facilitating the deployment of MER in practical applications. NaME is available at github.com/real-ljt/NAMEdataset.
The automatic generation of diverse and human-like facial reactions in dyadic dialogue remains a critical challenge for human-computer interaction systems. Existing methods fail to model the stochasticity and dynamics inherent in real human reactions. To address this, we propose ReactDiff, a novel temporal diffusion framework for generating diverse facial reactions that are appropriate for responding to any given dialogue context. Our key insight is that plausible human reactions demonstrate smoothness, and coherence over time, and conform to constraints imposed by human facial anatomy. To achieve this, ReactDiff incorporates two vital priors (spatio-temporal facial kinematics) into the diffusion process: i) temporal facial behavioral kinematics and ii) facial action unit dependencies. These two constraints guide the model toward realistic human reaction manifolds, avoiding visually unrealistic jitters, unstable transitions, unnatural expressions, and other artifacts. Extensive experiments on the REACT2024 dataset demonstrate that our approach not only achieves state-of-the-art reaction quality but also excels in diversity and reaction appropriateness. Our code is publicly available at https://github.com/lingjivoo/ReactDiff.
Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust Optimization
PDF ↗Dynamic Facial Expression Recognition (DFER) plays a critical role in affective computing and human-computer interaction. Although existing methods achieve comparable performance, they inevitably suffer from performance degradation under sample heterogeneity caused by multi-source data and individual expression variability. To address these challenges, we propose a novel framework, called Heterogeneity-aware Distributional Framework (HDF), and design two plug-and-play modules to enhance time-frequency modeling and mitigate optimization imbalance caused by hard samples. Specifically, the Time-Frequency Distributional Attention Module (DAM) captures both temporal consistency and frequency robustness through a dual-branch attention design, improving tolerance to sequence inconsistency and visual style shifts. Then, based on gradient sensitivity and information bottleneck principles, an adaptive optimization module Distribution-aware Scaling Module (DSM) is introduced to dynamically balance classification and contrastive losses, enabling more stable and discriminative representation learning. Extensive experiments on two widely used datasets, DFEW and FERV39k, demonstrate that HDF significantly improves both recognition accuracy and robustness. Our method achieves superior weighted average recall (WAR) and unweighted average recall (UAR) while maintaining strong generalization across diverse and imbalanced scenarios. Codes are released at https://github.com/QIcita/HDF_DFER.
Multi-modal emotion recognition has garnered increasing attention as it plays a significant role in human-computer interaction (HCI) in recent years. Since different discrete emotions may exist at the same time, compared with single-class emotion recognition, emotion distribution learning (EDL) that identifies a mixture of basic emotions has gradually emerged as a trend. However, existing EDL methods face challenges in mining the heterogeneity among multiple modalities. Besides, rich semantic correlations across arbitrary basic emotions are not fully exploited. In this paper, we propose a multi-modal emotion distribution learning framework, named HeLo, aimed at fully exploring the heterogeneity and complementary information in multi-modal emotional data and label correlation within mixed basic emotions. Specifically, we first adopt cross-attention to effectively fuse the physiological data. Then, an optimal transport (OT)-based heterogeneity mining module is devised to mine the interaction and heterogeneity between the physiological and behavioral representations. To facilitate label correlation learning, we introduce a learnable label embedding optimized by correlation matrix alignment. Finally, the learnable label embeddings and label correlation matrices are integrated with the multi-modal representations through a novel label correlation-driven cross-attention mechanism for accurate emotion distribution learning. Experimental results on two publicly available datasets demonstrate the superiority of our proposed method in emotion distribution learning.
Browser fingerprinting is a pervasive online tracking technique used increasingly often for profiling and targeted advertising. Prior research on the prevalence of fingerprinting heavily relied on automated web crawls, which inherently struggle to replicate the nuances of human-computer interactions. This raises concerns about the accuracy of current understandings of real-world fingerprinting deployments. As a result, this paper presents a user study involving 30 participants over 10 weeks, capturing telemetry data from real browsing sessions across 3,000 top-ranked websites. Our evaluation reveals that automated crawls miss almost half (45%) of the fingerprinting websites encountered by real users. This discrepancy mainly stems from the crawlers' inability to access authentication-protected pages, circumvent bot detection, and trigger fingerprinting scripts activated by specific user interactions. We also identify potential new fingerprinting vectors present in real user data but absent from automated crawls. Finally, we evaluate the effectiveness of federated learning for training browser fingerprinting detection models on real user data, yielding improved performance than models trained solely on automated crawl data.