Social Media Popularity Prediction is a complex multimodal task that requires effective integration of images, text, and structured information. However, current approaches suffer from inadequate visual-textual alignment and fail to capture the inherent cross-content correlations and hierarchical patterns in social media data. To overcome these limitations, we establish a multi-class framework , introducing hierarchical prototypes for structural enhancement and contrastive learning for improved vision-text alignment. Furthermore, we propose a feature-enhanced framework integrating dual-grained prompt learning and cross-modal attention mechanisms, achieving precise multimodal representation through fine-grained category modeling. Experimental results demonstrate state-of-the-art performance on benchmark metrics, establishing new reference standards for multimodal social media analysis.
论文检索
输入标题、作者或关键词,从 8,216 篇学术成果中精准定位
As embodied intelligence has become a new hot topic in current artificial intelligence research, facial reaction generation has increasingly become a key technology for achieving natural human-computer interaction. Existing methods typically rely on bimodal inputs of audio and visual signals, but they still suffer from poor cross-modal consistency and insufficient feature fusion, making it difficult to effectively capture complex facial features. To enhance the representational capacity of fused features, this paper proposes a Scattering-Conditioned Diffusion Model (SC-Diff), which extracts stable multi-scale structural features in the frequency domain via wavelet scattering transform and injects them into the diffusion generation process as conditional information, thereby enhancing the modeling ability of representing local facial variations. Furthermore, considering that different prediction tasks exhibit varying sensitivity to target changes during training, we introduce an uncertainty-based adaptive loss weighting strategy to dynamically balance three types of supervision targets: facial action units, facial affect, and facial expressions. Experimental results on the REACT 2025 dataset demonstrate that the proposed method outperforms existing state-of-the-art approaches across multiple evaluation metrics.
The rise of Multimodal Large Language Models (MLLMs) offers new opportunities for Micro-Expression (ME) analysis. This paper introduces Micro-Expression Visual Question Answering (ME-VQA), a novel task reformulating ME annotations (e.g., emotion categories, action units) into QA pairs. To address key challenges-hardware limitations, context inconsistency, and compositional reasoning gaps-we propose a Relationship-Aware Hierarchical VQA Framework. Our approach leverages mined emotion correlations (e.g., coarse-to-fine label dependencies) and employs a two-stage process: 1) Coarse-grained anchoring for broad emotion categories, and 2) Fine-grained reasoning constrained by coarse outputs and statistical rules. We further optimize efficiency via a dual-phase video sampling strategy: during training, keyframes (onset/apex/offset) and random non-expression frames are used; uniform sampling is applied at inference. Experiments demonstrate significant improvements in answer consistency and accuracy.
Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In recent years, substantial advancements have been made in the areas of ME recognition, spotting, and generation. However, conventional approaches that treat spotting and recognition as separate tasks are suboptimal, particularly for analyzing long-duration videos in realistic settings. Concurrently, the emergence of multimodal large language models (MLLMs) and large vision-language models (LVLMs) offers promising new avenues for enhancing ME analysis through their powerful multimodal reasoning capabilities. The ME grand challenge (MEGC) 2025 introduces two tasks that reflect these evolving research directions: (1) ME spot-then-recognize (ME-STR), which integrates ME spotting and subsequent recognition in a unified sequential pipeline; and (2) ME visual question answering (ME-VQA), which explores ME understanding through visual question answering, leveraging MLLMs or LVLMs to address diverse question types related to MEs. All participating algorithms are required to run on this test set and submit their results on a leaderboard. More details are available at https://megc2025.github.io.
Unmanned Aerial Vehicles, operating in environments with relatively few obstacles, offer high maneuverability and full three-dimensional mobility. This allows them to rapidly approach objects and perform a wide range of tasks often challenging for ground robots, making them ideal for exploration, inspection, aerial imaging, and everyday assistance. In this paper, we introduce AirStar, a UAV-centric embodied platform that turns a UAV into an intelligent aerial assistant: a large language model acts as the cognitive core for environmental understanding, contextual reasoning, and task planning. AirStar accepts natural interaction through voice commands and gestures, removing the need for a remote controller and significantly broadening its user base.It combines geospatial knowledge-driven long-distance navigation with contextual reasoning for fine-grained short-range control, resulting in an efficient and accurate vision-and-language navigation (VLN) capability. Furthermore, the system also offers built-in capabilities such as cross-modal question answering, intelligent filming, and target tracking. With a highly extensible framework, it supports seamless integration of new functionalities, paving the way toward a general-purpose, instruction-driven intelligent UAV agent.The supplementary PPT is available at https://buaa-colalab.github.io/airstar.github.io.
Text-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ''chain'' instructions where operations are interdependent. Current models struggle with these intricate directives, and existing benchmarks inadequately evaluate such capabilities. Specifically, they often overlook multi-instruction and chain-instruction complexities, and common consistency metrics are flawed. To address this, we introduce ComplexBench-Edit, a novel benchmark designed to systematically assess model performance on complex, multi-instruction, and chain-dependent image editing tasks. ComplexBench-Edit also features a new vision consistency evaluation method that accurately assesses non-modified regions by excluding edited areas. Furthermore, we propose a simple yet powerful Chain-of-Thought (CoT)-based approach that significantly enhances the ability of existing models to follow complex instructions. Our extensive experiments demonstrate ComplexBench-Edit's efficacy in differentiating model capabilities and highlight the superior performance of our CoT-based method in handling complex edits. The data and code are released at https://github.com/llllly26/ComplexBench-Edit.
Identifying individual animals within large wildlife populations is essential for effective wildlife monitoring and conservation efforts. Recent advancements in computer vision have shown promise in animal re-identification (Animal ReID) by leveraging data from camera traps. However, existing Animal ReID datasets rely exclusively on visual data, overlooking environmental metadata that ecologists have identified as highly correlated with animal behavior and identity, such as temperature and circadian rhythms. Moreover, the emergence of multimodal models capable of jointly processing visual and textual data presents new opportunities for Animal ReID, but existing datasets fail to leverage these models' text-processing capabilities, limiting their full potential. To address these limitations, we propose MetaWild, a multimodal Animal ReID dataset comprising 20,890 images across six species, paired with environmental metadata extracted from embedded camera trap overlays and scene contexts. Additionally, to facilitate the use of metadata in existing ReID methods, we propose the Meta-Feature Adapter (MFA), a lightweight module that can be incorporated into existing vision-language model (VLM)-based Animal ReID methods, allowing ReID models to leverage both environmental metadata and visual information to improve ReID performance. Experiments on MetaWild show that combining baseline ReID models with MFA to incorporate metadata consistently improves performance compared to using visual information alone, validating the effectiveness of incorporating metadata in re-identification. We hope that our proposed dataset can inspire further exploration of multimodal approaches for Animal ReID. Our dataset and supplementary materials are available at https://jim-lyz1024.github.io/MetaWild/.
The rapid advancements in Artificial Intelligence (AI) increasingly underscore the critical importance of multi-modal datasets for training robust and versatile models. However, the dominance of English-centric resources and the scarcity of datasets capturing diverse languages and cultural nuances limit AI's inclusivity and global applicability, posing a key challenge. Critically, language's intrinsic link to cultural context demands models possess profound cultural insight beyond mere translation to genuinely grasp societal norms and unique cultural expressions. Culturally rich broadcast content offers a solution; its systematic curation can mitigate linguistic/cultural imbalances, fostering culturally deep multi-modal datasets. This paper introduces the HAN (Korean Heritage Augmented Narrative Visual-Language Description) dataset, a new resource for multilingual image captioning and retrieval. HAN comprises 41,000 images captured from Korean broadcast video clips with 410,000 Korean/English narrative-style captions offering multifaceted perspectives on each visual instance. By incorporating Korean heritage, HAN reflects cultural diversity, addressing limitations of existing datasets focused on simple descriptions. Furthermore, this work analyzes HAN's caption diversity impact on retrieval, proposing strategies to enhance efficacy. These findings underscore HAN's potential to significantly advance multi-modal and multilingual processing, supported by its value as a rich resource for learning approaches that integrate diverse data types (like vision and text), natural language processing across various languages, and Korean heritage studies.
In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, with 6,035 masks and 91k frame-level 3D skeleton annotations. Additionally, we propose Cascaded Pose Perception (CPP), a benchmarking method that simultaneously performs language-driven mask segmentation and 3D pose estimation within a unified model. CPP first learns 2D pose information, utilizes a body fusion module to aid in mask segmentation, and employs a mask fusion module to mitigate mask noise before outputting 3D poses. Extensive evaluation of CPP and existing benchmarks on Panoptic-L3D demonstrates the necessity of this novel task and dataset for advancing 3DHPE. Our dataset is available at https://languagedriven3dposeestimation.github.io/.
Diffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and mask-and-inpainting. Among these, instruction-based editing stands out for its convenience and effectiveness in following human instructions across diverse scenarios. However, it still focuses on simple editing operations like adding, replacing, or deleting, and falls short of understanding aspects of world dynamics that convey the realistic dynamic nature in the physical world. Therefore, this work EditWorld introduces a new editing task, namely world-instructed image editing, which defines and categorizes the instructions grounded by various world scenarios. We curate a new image editing dataset with world instructions using a set of large pretrained models (e.g., GPT, Video-LLava and SDXL). To enable sufficient simulation of world dynamics for image editing, our EditWorld trains model in the curated dataset, and improves instruction-following ability with designed post-edit strategy. Extensive experiments demonstrate our method significantly outperforms existing editing methods in this new task. https://github.com/YangLing0818/EditWorld
Moving target selection in multimedia interactive systems faces unprecedented challenges as users increasingly interact across diverse, dynamic contexts-from live streaming in moving vehicles to VR gaming in varying environments. Existing approaches rely on probabilistic models that relate endpoint distribution to target properties (size, speed). However, these methods require substantial training data for each new context and lack transferability across scenarios, limiting their practical deployment in diverse multimedia environments where rich multimodal contextual information is readily available. This paper introduces MAGNeT (Multimodal Adaptive Gaussian Networks), which addresses these problems by combining classical statistical modeling with context-aware multimodal method. MAGNeT dynamically fuses pre-fitted Ternary-Gaussian models from various scenarios based on real-time contextual cues, enabling effective adaptation with minimal training data while preserving model interpretability. We take experiments on self-constructed 2D and 3D moving target selection datasets under in-vehicle vibration conditions. Extensive experiments demonstrate that MAGNeT achieves lower error rates with few-shot samples, by applying context-aware fusion of Gaussian experts from multi-factor conditions.
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
PDF ↗Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between temporal video sequences and textual semantics. Existing approaches, relying on static cross-attention or prompt-tuning mechanisms, fail to adaptively model the evolving relationships between modalities, leading to suboptimal alignment and limited generalization. Inspired by systems biology, we propose Reaction-Diffusion Multimodal Fusion (RDMF), a novel framework that reimagines video-language alignment as a reaction-diffusion (RD) process, drawing on the principles of pattern formation introduced by Alan Turing. In RDMF, video features diffuse across time to capture temporal context, while text-video interactions are modeled as non-linear reactions that amplify relevant features and suppress noise, forming emergent patterns akin to biological systems. Leveraging the Gray-Scott RD model, we design a computationally efficient fusion module that integrates video and text representations, supported by rigorous mathematical analysis of stability and convergence using Turing instability criteria. Our framework is theoretically grounded, employing advanced mathematical tools to ensure stable pattern formation, and is practically viable, incorporating standard components like pretrained encoders and DETR-style heads for moment retrieval and saliency prediction. RDMF represents a pioneering interdisciplinary approach, bridging systems biology and multimedia research to address the limitations of conventional multimodal fusion. Preliminary experiments demonstrate its potential to outperform existing methods in identifying salient video moments, offering a new paradigm for video-language tasks. Beyond its immediate applications, RDMF opens avenues for exploring biologically inspired architectures in multimedia, with implications for real-time video analysis, interactive media systems, and cross-disciplinary collaboration between multimedia and systems biology communities.
Longitudinal studies of animal vocalizations provide crucial insights into developmental patterns and communicative evolution. To aid such investigations in canines, this paper introduces the Canine Age Transition Vocalization Dataset, a large-scale collection of dog vocalizations featuring meticulously verified metadata (including precise birthdate, breed, and individual dog ID) for 125 dogs across 6 common breeds. Our in-depth longitudinal analysis of this dataset then reveals novel findings on how key vocal parameters, encompassing defined bark types and finer-grained acoustic components (Elemental Dog Bark Units, or EDBUs), change as dogs mature. This work, therefore, offers both a significant new resource and foundational data that enable deeper, more nuanced investigations into the lifelong vocal development of dogs and other animal communication.
Cold-start recommendation, which addresses the challenge of recommending items to new users without sufficient historical interaction data, remains difficult in personalized recommender systems. Cross-domain recommendation methods have gained attention in addressing this issue by transferring user preferences from a source domain to alleviate the cold-start problem in a target domain. Traditional embedding-based approaches typically rely on a one-to-one alignment over continuous user embeddings using overlapping user IDs, which often leads to severe overfitting and limited generalization, particularly when overlapping users between domains are sparse. In this paper, we propose a novel hierarchical recommendation framework specifically targeting the cold-start problem by modeling user preferences as hierarchical interests derived from textual embeddings and Residual Quantized Variational Autoencoders (RQ-VAE). Unlike traditional methods that directly align embeddings, our approach builds a mapping function transferring hierarchical user interest structures from the source to the target domain. This hierarchical mapping significantly enhances generalization, providing a more comprehensive and robust representation of user preferences. Additionally, we employ a generative rewriting mechanism utilizing Large Language Models to refine user-generated reviews into concise, semantically enriched summaries that explicitly highlight user interests. Extensive evaluations on Amazon review datasets demonstrate the effectiveness of our hierarchical preference modeling and generative rewriting approach, outperforming existing embedding-based methods consistently, especially in cold-start and sparsely populated scenarios. Our proposed method thus provides a robust, flexible solution for personalized recommendation tasks in cold-start conditions.
Foundation models have transformed multimedia analysis by enabling robust and transferable representations across diverse modalities and tasks. However, their static deployment conflicts with growing societal and regulatory demands-particularly the need to unlearn specific data upon request, as mandated by privacy frameworks such as the GDPR. Traditional unlearning approaches, including retraining, activation editing, or distillation, are often computationally expensive, fragile, and ill-suited for real-time or continuously evolving systems. In this paper, we propose a paradigm shift: rethinking unlearning not as a retroactive intervention but as a built-in capability. We introduce a prompt-based learning framework that unifies knowledge acquisition and removal within a single training phase. Rather than encoding information in model weights, our approach binds class-level semantics to dedicated prompt tokens. This design enables instant unlearning simply by removing the corresponding prompt-without retraining, model modification, or access to original data. Experiments demonstrate that our framework preserves predictive performance on retained classes while effectively erasing forgotten ones. Beyond utility, our method exhibits strong privacy and security guarantees: it is resistant to membership inference attacks, and prompt removal prevents any residual knowledge extraction, even under adversarial conditions. This ensures compliance with data protection principles and safeguards against unauthorized access to forgotten information, making the framework suitable for deployment in sensitive and regulated environments. Overall, by embedding removability into the architecture itself, this work establishes a new foundation for designing modular, scalable and ethically responsive AI models.
Concept-based models have been proposed as a new line of research for explainable by-design deep learning models. However, those models show their whole power when applied to benchmarks where the concepts are well defined and the concepts' attributes easily extractable from the raw data. In this paper, we challenge the most recent concept-based model initially developed for image classification, on more complex interpretative tasks from a recently proposed video benchmark where they perform poorly. We conduct a root cause analysis of the poor performances of state-of-the-art explainable concept-based models for these multimodal interpretative tasks, and propose adaptations to design robust explainable models for detecting character objectification in this novel challenging video benchmark. We show that the optimal architectural choice may vary depending on the modality setting, thereby showing that designing multimodal concept-based approaches remains an open challenge and calls for further investigation.
The dynamic variations of food quality across spatial and temporal scales pose significant challenges for global food safety and nutrition research, requiring comprehensive analysis of diverse, multi-modal, and distributed data while preserving privacy. Existing centralized approaches suffer from data silos and limited collaboration, and although federated learning and blockchain technologies have shown promise independently, their combined potential for incentivized, privacy-preserving, and heterogeneous model collaboration remains underexplored. In this paper, we propose the concept of a Global Spatial-Temporal Food Memory-a novel research paradigm that envisions secure, decentralized, and incentivized collaboration among multiple stakeholders worldwide, leveraging blockchain-enabled token-based rewards integrated with federated learning of heterogeneous models. We discuss the scientific challenges and opportunities inherent in this vision, including multi-modal data fusion, trustworthy incentive mechanisms, and scalable long-term temporal analysis. This work aims to open new avenues in multimedia research by bridging decentralized AI, blockchain, and spatiotemporal food quality monitoring, providing a foundation for future explorations in privacy-preserving, collaborative, and large-scale multimedia data analysis.
Commonaiverse is an interactive installation exploring human emotions through full-body motion tracking and real-time AI feedback. Participants engage in three phases: Teaching, Exploration and the Cosmos Phase, collaboratively expressing and interpreting emotions with the system. The installation integrates MoveNet for precise motion tracking and a multi-recommender AI system to analyze emotional states dynamically, responding with adaptive audiovisual outputs. By shifting from top-down emotion classification to participant-driven, culturally diverse definitions, we highlight new pathways for inclusive, ethical affective computing. We discuss how this collaborative, out-of-the-box approach pushes multimedia research beyond single-user facial analysis toward a more embodied, co-created paradigm of emotional AI. Furthermore, we reflect on how this reimagined framework fosters user agency, reduces bias, and opens avenues for advanced interactive applications.
Low-light image enhancement (LLIE) aims to restore low-light images to normal lighting conditions by improving their illumination and fine details, thereby facilitating efficient execution of downstream visual tasks. Traditional LLIE methods improve image quality but often introduce high-frequency artifacts, which are difficult to eliminate, hindering detail recovery and quality enhancement in LLIE. To solve this problem, we introduce a novel perspective: instead of traditional artifact suppression, sparsification-induced artifacts are repurposed as constructive regularization signals to guide detail recovery. By analyzing the impact of sparsified frequency components and their role in reconstruction artifacts, a detailed mathematical framework is presented. Specifically, we propose a novel loss function SASW-Loss which combining Sparse Artifact Similarity Loss (SAS-Loss) and Walsh-Hadamard Coefficient Loss (WHC-Loss). SAS-Loss mitigates the over-compensation of missing frequencies, helping the network recover structural details, while WHC-Loss optimizes the frequency-domain representation, restoring luminance, suppressing noise, and enhancing both structure and details. Extensive experiments show that our approach outperforms existing state-of-the-art methods, achieving superior performance in structural detail preservation and noise suppression. These results validate the effectiveness of our new perspective, which leverages sparsification artifacts to guide detail recovery, demonstrating significant improvements and robust performance across multiple models, and opening new avenues for future research. The code is available at https://github.com/werringwu/SASW.git.
Fine-grained cross-view localization seeks to predict ground-level camera positions within GPS-tagged aerial images by matching ground and aerial views. Existing methods often rely on large-scale ground truth annotations from specific regions, but performance degrades due to domain shifts when models trained in one area are applied to another. However, collecting region-specific annotations for each area is costly or infeasible. To address this, we propose a self-distillation curriculum learning framework that generalizes pretrained localization models to unseen new areas. Our approach introduces a Dirichlet-based quality assessment strategy to evaluate teacher-generated pseudo labels, where high uncertainty signals noisy predictions and low uncertainty indicates clean samples. This uncertainty is used to guide an easy-to-hard curriculum learning strategy, where easy samples are prioritized initially, and more challenging samples are progressively incorporated, enabling effective student training. Furthermore, we develop a joint optimization scheme that updates both the student model and pseudo labels, applying adaptive label smoothing to mitigate label noises and taking full advantage of new area data. Extensive experimental results on the VIGOR and KITTI benchmarks demonstrate that our method outperforms state-of-the-art approaches in new area localization, achieving superior accuracy without additional supervision.