Learning natural and diverse behaviors from human motion datasets remains a significant challenge in physics-based character control. Existing conditional adversarial models often suffer from tight and biased embedding distributions where embeddings from the same motion are closely grouped in a small area, and shorter motions occupy even less space. Our empirical observations indicate this limits the representational capacity and diversity under each skill. An ideal latent space should be maximally packed by all motion's embedding clusters. Although methods that employ separate embedding spaces for each motion mitigate this limitation to some extent, introducing a hybrid discrete-continuous embedding space imposes a huge exploration burden on the high-level policy. To address the above limitations, we propose a versatile skill-conditioned controller that learns diverse skills with expressive variations. Our approach leverages the Neural Collapse phenomenon, a natural outcome of the classification-based encoder, to uniformly distribute cluster centers. We additionally propose a novel Embedding Expansion technique to form stylistic embedding clusters for diverse skills that are uniformly distributed on a hypersphere, maximizing the representational area occupied by each skill and minimizing unmapped regions. This maximally packed and uniformly distributed embedding space ensures that embeddings within the same cluster generate behaviors conforming to the characteristics of the corresponding motion clips, yet exhibiting noticeable variations within each cluster. Compared to existing methods, experimental results demonstrate that our controller not only generates high-quality, diverse motions covering the entire dataset but also achieves superior controllability, motion coverage, and diversity under each skill. Both qualitative and quantitative results confirm these traits, enabling our controller to be applied to a wide range of downstream tasks and serving as a cornerstone for diverse applications.
论文检索
输入标题、作者或关键词,从 7,876 篇学术成果中精准定位
High-quality thermal facial data is essential for advancing biometric recognition, surveillance, in-cabin driver monitoring, and human-computer interaction, all of which are integral for modern multimedia and interactive AI systems. In this work, we optimized the FLUX text-to-image diffusion model on diverse real-world thermal facial datasets to generate hyper-realistic 2D thermal facial images for both males and females, and propose a new dataset, ThermVision. To enhance their multimedia applicability, these images are processed through a video retargeting pipeline, where driving videos animate realistic facial expressions and head pose variations from a single 2D thermal image, producing high-fidelity thermal facial video sequences. The overall rendered dataset incorporates smart transformations, ensuring diversity across gender balance, extreme head pose variations, expressive facial dynamics, and facial accessories, making it a valuable resource for real-world applications. Additionally, we provide facial detection annotations to facilitate precise feature extraction and thermal-face analysis. To validate our synthetic dataset, we evaluate its effectiveness in thermal gender classification, as downstream machine learning task, along with thermal face localization and facial landmarks detection demonstrating its applicability in real-world scenarios. This approach significantly improves the availability, realism, and integration of thermal facial data, paving the way for more robust and immersive AI-powered thermal imaging applications. The dataset, code and associated models are available at- https://mali-farooq.github.io/ThermVision/
Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker identity. We identify three critical limitations in current emotional talking head generation: insufficient utilization of audio's inherent emotional cues, identity leakage in emotion representations, and isolated learning of emotion correlations. To address these challenges, we propose a novel framework dubbed as DICE-Talk, following the idea of disentangling identity with emotion, and then cooperating emotions with similar characteristics. First, we develop a disentangled emotion embedder that jointly models audio-visual emotional cues through cross-modal attention, representing emotions as identity-agnostic Gaussian distributions. Second, we introduce a correlation-enhanced emotion conditioning module with learnable emotion banks that explicitly capture inter-emotion relationships through vector quantization and attention-based feature aggregation. Third, we design an emotion discrimination objective that enforces affective consistency during the diffusion process through latent-space classification. Extensive experiments on MEAD and HDTF datasets demonstrate our method's superiority, outperforming state-of-the-art approaches in emotion accuracy while maintaining competitive lip-sync performance. Qualitative results and user studies further confirm our method's ability to generate identity-preserving portraits with rich, correlated emotional expressions that naturally adapt to unseen identities.
We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining ) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector ). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.
End-to-end automated fact-checking (AFC) aims to assess the truthfulness of claims using retrieved evidence. Some researchers use crawlers or search APIs to retrieve evidence from the web for veracity classification. However, existing methods indiscriminately rely on the retrieved evidence and overlook that the retrieved results are not always reliable. This unilateral reliance on evidence significantly hampers the performance of fact-checking. In this paper, we account for the diverse reliability levels of retrieved evidence and eliminate the negative impact from the causal perspective. To achieve our goal, we propose a novel Causal intervention and Counterfactual reasoning based Multi-Checker framework (CCMC), which introduces two additional counterfactual fact-checkers to verify claims from the counterfactual perspective. Specifically, we construct two distinct types of counterfactual instances via causal intervention to imitate the situation where the evidence is partially reliable or totally unreliable. Correspondingly, two counterfactual fact-checkers are trained with tailored counterfactual instances by counterfactual reasoning. During inference, the two counterfactual fact-checkers are employed to estimate and eliminate the potential impact of unreliable evidence. Extensive experiments on two real-world datasets demonstrate the superiority of our approach for improving end-to-end AFC. Especially, we surpass existing methods by 3.70% and 5.55% under gold and system evidence on the MOCHEG benchmark, respectively. Our code is available at https://github.com/BeiyuXuboL/CCMC.
Remote sensing image classification with noisy labels is receiving increasing attention. However, the existing methods ignore the context information of the training sample and judge whether the label is a noise label only by monitoring the loss value of a single sample, which may lead to misjudgment of the sample label. Additionally, these algorithms do not consider constructing pairs of confidence instances to obtain robust potential representations after identifying confidence instances. In this paper, a Multi-view Collaborative Representation Learning (MCRL) approach from noisy labels is proposed to improve the classification performance of very high resolution (VHR) remote sensing images. Specifically, we design a correction strategy based on spatial consistency and confidence-aware mechanisms. This strategy quantitatively measures label reliability by mining the contextual information of labelled samples within the adaptive region. Leveraging the spatial consistency principle and the confidence-aware mechanism to correct and smooth the noisy labels progressively. Moreover, we construct confidence sample pairs by establishing relationships between samples within and between views to obtain robust latent representations, which improves the model's tolerance to noisy labels. Experiments show that the MCRL can significantly reduce the impact of noisy labels on the model and is more competitive than homologous algorithms.
Medical time series, such as Electroencephalogram (EEG) and Electrocardiogram (ECG), are widely used for disease detection, with multiple electrodes or sensors recording simultaneously. Accurately modeling inter-channel relationships is crucial for improving detection performance. Current methods mainly rely on data-driven approaches to model channel relationships, facing two challenges: (1) insufficient integration of medical prior knowledge, hindering the accurate representation of physiological correlations between channels, and (2) high temporal pattern similarity across channels, leading to feature redundancy and degraded classification performance. To address these issues, we introduce KEMed, a knowledge-augmented model for medical time series classification. The model incorporates medical textual prior knowledge by generating natural language descriptions for each channel and leveraging Pre-trained Language Model (PLM) for semantic representation, enabling precise identification of physiological and pathological similarities and differences between channels. Specifically, KEMed optimizes channel relationships through knowledge-guided clustering and weighting mechanisms and leverages Large Language Model (LLM) to capture spatiotemporal dependencies, thereby enhancing classification performance. Experimental results on five medical time series datasets demonstrate that KEMed consistently outperforms state-of-the-art methods, validating the effectiveness and superiority of knowledge augmentation in medical time series classification.
Graph classification is a fundamental machine learning problem with extensive applications in multimedia and biochemical analysis. Contemporary graph classification models usually require precise graph labels for supervision, even after self-supervised pre-training. However, in practical applications, the extensive precise annotation of graphs could be expensive or impractical. To exploit data efficiently, this work studies partial label graph learning, in which each graph is linked to a set of candidate labels but only one of them is accurate. Label ambiguity would bring difficulties in extracting graph semantics and the risk of overfitting noisy partial labels. Here, we present a novel approach called Coupled Dual Separation (CODE). To improve graph semantics mining under label ambiguity, our CODE contains a message passing branch and a graph kernel branch, which explore graph semantics implicitly and explicitly, respectively. To facilitate information exchange, we utilize one branch to separate partially labeled graphs into an informative set and an uninformative set, which provides guidance for the optimization of the other branch. Furthermore, to mitigate the risk of overfitting, parameters in coupled branches are partitioned into critical and non-critical ones for separated optimization procedures. Extensive experiments on several benchmark datasets validate the effectiveness of the proposed CODE.
Existing prototype learning-based Multiple Instance Learning (MIL) methods mainly focus on learning a single set of prototypes for each class or generating a generic prototype from the overall data distribution. This design forces the model to compress general and heterogeneous features into identical prototype embeddings, prioritizing general features over subtle but discriminative features. Additionally, these methods often guide prototype updates by jointly optimizing attention score distributions and the distances between instances and prototypes, resulting in prototype biases due to over-concentration. To address these issues, we propose a dual prototype learning MIL (DP-MIL) framework that introduces two distinct sets of prototypes: primary prototypes, which capture general WSI features, and boundary prototypes, which capture discriminative features near the decision boundary. The DP-MIL framework employs three prototype-tailored losses: an alienation loss to encourage primary prototypes to be distant from decision boundaries, an affinity loss to anchor boundary prototypes near these boundaries, and a distance loss to enforce separation between the two prototype sets. To mitigate prototype semantic drift during training, we introduce a prototype joint updating and refinement strategy: for each prototype, we use its corresponding global token to filter out the most similar instances to momentum update the corresponding prototype set, while the boundary prototype set is refined with the mean pooled feature of hard samples. Extensive experiments on four datasets demonstrate the effectiveness of our DP-MIL framework and prototype updating strategy.
Image steganalysis is a detection task to distinguish whether a secret message is embedded in a digital image. Due to the domain inconsistency caused by Cover Source Mismatch(CSM) and Steganographic Algorithm Mismatch (SAM), most of them suffer from significant performance degradation. Recent mismatched steganalysis focused on extracting domain invariant features by domain adversarial training or feature alignment. However these schemes are limited to unstable performance in diverse domain mismatch scenarios, and are even ineffective in some cases. In this paper, we propose a Universal Mismatched Steganalysis PCD-UMS via pair-wise confidence difference-based pseudo-label selection from the perspective of optimizing target training data. Specifically, we reveal a strong positive correlation commonality between pair-wise confidence difference and the detection performance of steganalysis among various mismatch scenarios. Based on this, a novel pseudo-label selection strategy consisting of maximum confidence difference first (MCDF) rule and pair-wise label differential storage (PLDS) rule is designed to select and filter the reliable target pseudo-labels. Furthermore, a multi-perspective pair-wise feature alignment loss is designed to initially transfer the classification ability of source steganalysis, thus solving the problem that source steganalysis fails completely under some domain mismatch scenarios. Comprehensive experiments show that our PCD-UMS outperforms the existing mismatched steganalysis by 12.07% and 3.40% in terms of detection performance under CSM and SAM scenarios.
Galaxy morphology analysis involves studying galaxies based on their shapes and structures. For such studies, fundamental tasks include identifying and classifying galaxies in astronomical images, as well as retrieving visually or structurally similar galaxies through similarity search. Existing methods either directly train domain-specific foundation models on large, annotated datasets or fine-tune vision foundation models on a smaller set of images. The former is effective but costly, while the latter is more resource-efficient but often yields lower accuracy. To address these challenges, we introduce GalaxAlign, a multimodal approach inspired by how citizen scientists identify galaxies in astronomical images by following textual descriptions and matching schematic symbols. Specifically, GalaxAlign employs a tri-modal alignment framework to align three types of data during fine-tuning: (1) schematic symbols representing galaxy shapes and structures, (2) textual labels for these symbols, and (3) galaxy images. By incorporating multimodal instructions, GalaxAlign eliminates the need for expensive pretraining and enhances the effectiveness of fine-tuning. Experiments on galaxy classification and similarity search demonstrate that our method effectively fine-tunes general pre-trained models for astronomical tasks by incorporating domain-specific multi-modal knowledge. Code is available at https://github.com/RapidsAtHKUST/GalaxAlign.
Electroencephalogram (EEG) signal classification faces significant challenges due to data distribution shifts caused by heterogeneous electrode configurations, acquisition protocols, and hardware discrepancies across domains. This paper introduces IMAC, a novel channel-dependent mask and imputation self-supervised framework that formulates the alignment of cross-domain EEG data shifts as a spatial time series imputation task. To address heterogeneous electrode configurations in cross-domain scenarios, IMAC first standardizes different electrode layouts using a 3D-to-2D positional unification mapping strategy, establishing unified spatial representations. Unlike previous mask-based self-supervised representation learning methods, IMAC introduces spatio-temporal signal alignment. This involves constructing a channel-dependent mask and reconstruction task framed as a low-to-high resolution EEG spatial imputation problem. Consequently, this approach simulates cross-domain variations such as channel omissions and temporal instabilities, thus enabling the model to leverage the proposed imputer for robust signal alignment during inference. Furthermore, IMAC incorporates a disentangled structure that separately models the temporal and spatial information of the EEG signals separately, reducing computational complexity while enhancing flexibility and adaptability. Comprehensive evaluations across 10 publicly available EEG datasets demonstrate IMAC's superior performance, achieving state-of-the-art classification accuracy in both cross-subject and cross-center validation scenarios. Notably, IMAC shows strong robustness under both simulated and real-world distribution shifts, surpassing baseline methods by up to 35% in integrity scores while maintaining consistent classification accuracy.
Early screening for Alzheimer's Disease (AD) through speech presents a promising non-invasive approach. However, challenges such as limited data and the lack of fine-grained, adaptive feature selection often hinder performance. To address these issues, we propose MoTAS, a robust framework designed to enhance AD screening efficiency. MoTAS leverages Text-to-Speech (TTS) augmentation to increase data volume and employs a Mixture of Experts (MoE) mechanism to improve multimodal feature selection, jointly enhancing model generalization. The process begins with automatic speech recognition (ASR) to obtain accurate transcriptions. TTS is then used to synthesize speech that enriches the dataset. After extracting acoustic and text embeddings, the MoE mechanism dynamically selects the most informative features, optimizing feature fusion for improved classification. Evaluated on the ADReSSo dataset, MoTAS achieves a leading accuracy of 85.71%, outperforming existing baselines. Ablation studies further validate the individual contributions of TTS augmentation and MoE in boosting classification performance. These findings highlight the practical value of MoTAS in real-world AD screening scenarios, particularly in data-limited settings.
Forged videos are often subjected to double compression. When a forger maliciously or unintentionally increases the video's bitrate during re-encoding, the resulting videos are termed fake bitrate videos. Detecting these videos offers a generalized approach for efficiently identifying potentially forged content within large datasets. However, previous research has largely focused on video-level detection of fully fake bitrate videos, where an entire video is re-encoded at a higher bitrate after content modification or the creation of fake high-definition (HD) footage. In practice, a skilled forger may adjust the bitrate of only specific video segments, generating partial fake bitrate videos-a common manipulation in tampering processes like video splicing. Existing methods face difficulties in detecting such partial modifications at the frame level and in pinpointing the manipulated segments. Our study addresses this gap by introducing a novel frame-level detection approach, which significantly enhances forensic precision. We simultaneously account for two types of abnormal frames arising from re-encoding and bitrate escalation and, for the first time, define fake bitrate video detection as a triple classification problem. To meet the challenges of this task, we extract anomalous bitrate-compression traces that capture subtle differences among the three frame types. Additionally, we propose the Trident Transformer Network (TTNet), a model designed to effectively integrate and learn high-frequency information within the encoding domain. Our approach achieves substantial improvements in accuracy, surpassing state-of-the-art methods by 3.62% and 11.95% in video-level and frame-level detection scenarios, respectively.
The integration of prompt tuning with multimodal learning has shown significant generalization abilities for various downstream tasks. Despite advancements, existing methods heavily depend on massive modality-specific labeled data (e.g., video, audio, and image), or are customized for a single modality. In this study, we present Text as Any-Modality by Consistent Prompt Tuning (TaAM-CPT), a scalable approach for constructing a general representation model toward unlimited modalities using solely text data. TaAM-CPT comprises modality prompt pools, text construction, and modality-aligned text encoders from pre-trained models, which allows for extending new modalities by simply adding prompt pools and modality-aligned text encoders. To harmonize the learning across different modalities, TaAM-CPT designs intra- and inter-modal learning objectives, which can capture category details within modalities while maintaining semantic consistency across different modalities. Benefiting from its scalable architecture and pre-trained models, TaAM-CPT can be seamlessly extended to accommodate unlimited modalities. Remarkably, without any modality-specific labeled data, TaAM-CPT achieves leading results on diverse datasets spanning various modalities, including video classification, image classification, and audio classification. The code is available at https://github.com/Jinx630/TaAM-CPT.
Accurate prediction of placental diseases via whole slide images (WSIs) is critical for preventing severe maternal and fetal complications. However, WSI analysis presents significant computational challenges due to the massive data volume. Existing WSI classification methods encounter critical limitations: (1) inadequate patch selection strategies that either compromise performance or fail to sufficiently reduce computational demands, and (2) the loss of global histological context resulting from patch-level processing approaches. To address these challenges, we propose an Efficient multimodal framework for Patient-level placental disease Diagnosis, named EmmPD. Our approach introduces a two-stage patch selection module that combines parameter-free and learnable compression strategies, optimally balancing computational efficiency with critical feature preservation. Additionally, we develop a hybrid multimodal fusion module that leverages adaptive graph learning to enhance pathological feature representation and incorporates textual medical reports to enrich global contextual understanding. Extensive experiments conducted on both a self-constructed patient-level Placental dataset and two public datasets demonstrating that our method achieves state-of-the-art diagnostic performance. The code is available at https://github.com/ECNU-MultiDimLab/EmmPD.
Sketching is a quick ideation and multimedia tool for effectively expressing design intent. By translating simple strokes into CAD models, it allows non-expert users to create editable designs, reducing the learning curve associated with traditional CAD software. However, current sketch-based CAD modeling methods are often limited to basic shapes and require structured inputs, making them less robust when dealing with varied sketch styles. To overcome these challenges, we propose a novel sketch-based modeling framework DAFU-CAD, that is both efficient and robust. Our approach features a Depth-Assisted and Feature-Unraveling sketch classification module that categorizes sketches into corresponding modeling operations, independent of their drawing style. A parameter regression and optimization module then estimates the modeling parameters, ensuring consistent and stable model reconstruction across different sketch inputs. To support this, we compile a diverse sketch dataset with a range of modeling categories and abstraction levels. Experimental results show that our method outperforms existing approaches in terms of both robustness and versatility.
Unified Medical Image Segmentation (UMIS) is critical for comprehensive anatomical assessment but faces challenges due to multi-scale structural heterogeneity. Conventional pixel-based approaches, lacking object-level anatomical insight and inter-organ relational modeling, struggle with morphological complexity and feature conflicts, limiting their efficacy in UMIS. We propose Mamba Snake, a novel deep snake framework enhanced by state space modeling for UMIS. Mamba Snake frames multi-contour evolution as a hierarchical state space atlas, effectively modeling macroscopic inter-organ topological relationships and microscopic contour refinements. We introduce a snake-specific vision state space module, the Mamba Evolution Block (MEB), which leverages effective spatiotemporal information aggregation for adaptive refinement of complex morphologies. Energy map shape priors further ensures robust long-range contour evolution in heterogeneous data. Additionally, a dual-classification synergy mechanism is incorporated to concurrently optimize detection and segmentation, mitigating under-segmentation of microstructures in UMIS. Extensive evaluations across five clinical datasets reveal Mamba Snake's superior performance.
Generative models have emerged as powerful tools capable of generating photorealistic images, spawning a wide range of applications across various domains. However, effectively integrating generative models into image classification tasks remains an open problem. Our analysis reveals that current generative data augmentation methods, as well as traditional data augmentation techniques, have limitations in simultaneously ensuring both fidelity (faithful foreground) and diversity (rich background contexts). To address this challenge, we propose Decomposition-Recomposition Data Augmentation (DRMix), an innovative intra-class data augmentation method. DRMix decomposes images into foreground-background and foreground parts, then performs diversified background recomposition and intra-class foreground recomposition, achieving dual diversity enhancement at both the image and part levels, and strikes a better trade-off between fidelity and diversity. Experimental results demonstrate that DRMix significantly improves performance across multiple tasks, including image classification, few-shot learning, and weakly-supervised object localization (WSOL).
Federated graph classification has emerged as a promising paradigm for privacy-preserving graph learning across distributed clients. However, real-world federated scenarios often suffer from severe data heterogeneity and label noise, which significantly degrade model performance. To address these challenges, we propose FedRog, a robust and personalized federated graph neural network framework that improves generalization under non-IID and noisy label settings. FedRog introduces a parameter-aware selection and fine-tuning mechanism to align global and local representations, and a neighbor embedding consistency constraint to enhance robustness against noisy supervision. Furthermore, a fine-grained, importance-guided global aggregation strategy based on Fisher information is employed to mitigate unreliable updates from low-quality clients. We conduct extensive experiments on 16 graph classification datasets under five heterogeneous data partition settings. Results show that FedRog consistently achieves competitive or superior performance compared to 14 baselines in terms of both accuracy and robustness under clean and noisy conditions.