论文检索

输入标题、作者或关键词,从 1,620 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
1,620篇论文
第 74 / 81 页

Hongming Wang, Yifeng Wu, Huimin Huang 0002, Hongtao Wu, Jiaxuan Jiang 0001, Xiaodong Zhang, Hao Zheng 0008, Yawen Huang, Xian Wu 0001, Yefeng Zheng 0001 等

The segmentation of substantial brain lesions is a significant and challenging task in the field of medical image segmentation. Substantial brain lesions in brain imaging exhibit high heterogeneity, with indistinct boundaries between lesion regions and normal brain tissue. Small lesions in single slices are difficult to identify, making the accurate and reproducible segmentation of abnormal regions, as well as their feature description, highly complex. Existing methods have the following limitations: 1) They rely solely on single-modal information for learning, neglecting the multi-modal information commonly used in diagnosis. This hampers the ability to comprehensively acquire brain lesion information from multiple perspectives and prevents the effective integration and utilization of multi-modal data inputs, thereby limiting a holistic understanding of lesions. 2) They are constrained by the amount of data available, leading to low sensitivity to small lesions and difficulty in detecting subtle pathological changes. 3) Current SAM-based models rely on external prompts, which cannot achieve automatic segmentation and, to some extent, affect diagnostic efficiency.To address these issues, we have developed a large-scale fully automated segmentation model specifically designed for brain lesion segmentation, named BrainSegDMIF. This model has the following features: 1) Dynamic Modal Interactive Fusion (DMIF) module that processes and integrates multi-modal data during the encoding process, providing the SAM encoder with more comprehensive modal information. 2) Layer-by-Layer Upsampling Decoder, enabling the model to extract rich low-level and high-level features even with limited data, thereby detecting the presence of small lesions. 3) Automatic segmentation masks, allowing the model to generate lesion masks automatically without requiring manual prompts.We tested and evaluated our model on two common brain disease segmentation benchmarks, including cases of focal cortical dysplasia and gliomas. Our model outperformed existing state-of-the-art methods across four metrics.

Mingle Zhou, Jiahui Liu, Jin Wan, Gang Li 0005, Min Li 0033

Unsupervised Continuous Anomaly Detection (UCAD) is gaining attention for effectively addressing the catastrophic forgetting and heavy computational burden issues in traditional Unsupervised Anomaly Detection (UAD). However, existing UCAD approaches that rely solely on visual information are insufficient to capture the manifold of normality in complex scenes, thereby impeding further gains in anomaly detection accuracy. To overcome this limitation, we propose an unsupervised continual anomaly detection framework grounded in multimodal prompting. Specifically, we introduce a Continual Multimodal Prompt Memory Bank (CMPMB) that progressively distills and retains prototypical normal patterns from both visual and textual domains across consecutive tasks, yielding a richer representation of normality. Furthermore, we devise a Defect-Semantic-Guided Adaptive Fusion Mechanism (DSG-AFM) that integrates an Adaptive Normalization Module (ANM) with a Dynamic Fusion Strategy (DFS) to jointly enhance detection accuracy and adversarial robustness. Benchmark experiments on MVTec AD and VisA datasets show that our approach achieves state-of-the-art (SOTA) performance on image-level AUROC and pixel-level AUPR metrics.

Bo Xu 0023, Jie Wei, Hongya Wang, Ming Du 0002, Hui Song, Yanghua Xiao

Multimodal Named Entity Recognition (MNER) integrates visual information to resolve textual ambiguities but struggles with generalizing to unseen entities (out-of-vocabulary, OOV), particularly in social media. To bridge this gap, we leveraging internal label knowledge and visual information and propose a Label-Enhanced Information Bottleneck Distillation (LIBD) framework, which transfers label-aware generalization capabilities via a teacher-student architecture. Our method introduces Dual-level Label Augmentation (DLA), enhancing the teacher model by integrating word-level entity replacement with labels and embedding-level learnable label vectors. This is paired with Information Bottleneck Distillation (IBD), selectively distilling critical knowledge from the teacher while suppressing irrelevant noise. Experiments on benchmark datasets demonstrate that LIBD outperforms state-of-the-art methods, especially in identifying OOV entities.

Yao Zhang, Ping Huang, Rui Zhang

The integration of Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) offers a new paradigm for complex decision-making tasks, especially through methods that optimize actor populations and critics collaboratively (e.g., DBCEM-TD3, QD-RL), collectively referred to as the Actor-Population Critic (APC) framework. However, existing methods still face two major challenges: traditional feature inputs fail to fully capture the spatial relationships, and the design of the critic struggles to balance both the quality and diversity of the policies. To address these issues, we propose Multimodal Dual Population Evolutionary Reinforcement Learning (M-DPERL), which achieves breakthroughs through cross-modal feature augmentation and dual-population coevolution. On one hand, a feature-image bimodal input enhancement mechanism is proposed, which dynamically encodes environmental features into spatial heatmaps. On the other hand, this method introduces a critic population into APC and, for the first time, proposes a population-guided fitness metric to optimize the critic's ability to guide the actor population in balancing quality and diversity. Additionally, we design the Flow-Fix Dynamics (FFD) mechanism to regulate the update rhythm of the dual populations and alleviate the evolutionary chaos in their coevolution. The results across a series of MUJOCO tasks demonstrate that M-DPERL significantly outperforms the baselines, with a 19.2% improvement in sample efficiency and a 17.1% increase in final performance.

Zhenxi Wang, Zongyao Yin, Yujie Hou, Xianchuan Yu

Recently, contrastive learning has emerged as a promising approach for multi-view clustering (MVC), as it enforces cross-view consistency and leverages complementary information from different views to enhance the analysis of heterogeneous data. However, traditional contrastive MVC methods suffer from an inherent limitation: their one-to-many contrast mechanism induces the False Negative Problem (FNP), where semantically similar intra-class instances are erroneously repelled. This phenomenon compromises intra-class consistency and ultimately degrades clustering performance. To overcome this issue, we propose a novel Pseudo lA bel gU ided univerS um lE arning (PAUSE) framework for robust multi-view clustering. Specifically, PAUSE operates in two synergistic stages: (1) A warm-up stage that employs dual contrastive learning to generate reliable pseudo-labels, establishing robust semantic relationships; (2) A fine-tuning stage that synthesizes universum samples via Mixup between anchor instances and out-of-class centroids, guided by the acquired pseudo-labels. This unique mechanism constructs generalized negative classes that expand inter-class margins while preserving intra-class cohesion. Crucially, the widened decision boundaries prevent misclassification of displaced intra-class instances, effectively circumventing FNP without requiring explicit negative pair correction. We further devise a robust universum contrastive loss that explicitly enforces cross-view consistency through adaptive boundary constraints. Extensive experiments on five multi-view benchmarks demonstrate that our PAUSE consistently outperforms 11 state-of-the-art multi-view learning methods. Our code is accessible at: https://github.com/xixi-555/PAUSE_main_code.

Peirong Zhang 0001, Kai Ding 0009, Lianwen Jin

In this paper, we propose SPECTRUM, a temporal-frequency synergistic model that unlocks the untapped potential of multi-domain representation learning for online handwriting verification (OHV). SPECTRUM comprises three core components: (1) a multi-scale interactor that finely combines temporal and frequency features through dual-modal sequence interaction and multi-scale aggregation, (2) a self-gated fusion module that dynamically integrates global temporal and frequency features via self-driven balancing. These two components work synergistically to achieve micro-to-macro spectral-temporal integration. (3) A multi-domain distance-based verifier then utilizes both temporal and frequency representations to improve discrimination between genuine and forged handwriting, surpassing conventional temporal-only approaches. Extensive experiments demonstrate SPECTRUM's superior performance over existing OHV methods, underscoring the effectiveness of temporal-frequency multi-domain learning. Furthermore, we reveal that incorporating multiple handwritten biometrics fundamentally enhances the discriminative power of handwriting representations and facilitates verification. These findings not only validate the efficacy of multi-domain learning in OHV but also pave the way for future research in multi-domain approaches across both feature and biometric domains. Code is publicly available at https://github.com/NiceRingNode/SPECTRUM.

Jiahuan Long, Wen Yao 0001, Tingsong Jiang, Jiacheng Hou, Shuai Jia, Junqi Wu 0002, Xiaoya Zhang, Xiaohu Zheng, Chao Ma 0004

Adversarial patches are widely used to evaluate the robustness of object detection systems in real-world scenarios. These patches were initially designed to deceive single-modal detectors (e.g., visible or infrared) and have recently been extended to target visible-infrared dual-modal detectors. However, existing dual-modal adversarial patch attacks have limited attack effectiveness across diverse physical scenarios. To address this, we propose CDUPatch, a universal cross-modal patch attack against visible-infrared object detectors across scales, views, and scenarios. Specifically, we observe that color variations lead to different levels of thermal absorption, resulting in temperature differences in infrared imaging. Leveraging this property, we propose an RGB-to-infrared adapter that maps RGB patches to infrared patches, enabling unified optimization of cross-modal patches. By learning an optimal color distribution on the adversarial patch, we can manipulate its thermal response and generate an adversarial infrared texture. Additionally, we introduce a multi-scale clipping strategy and construct a new visible-infrared dataset, MSDrone, which contains aerial vehicle images in varying scales and perspectives. These data augmentation strategies enhance the robustness of our patch in real-world conditions. Experiments on four benchmark datasets (e.g., DroneVehicle, LLVIP, VisDrone, MSDrone) show that our method outperforms existing patch attacks in the digital domain. Extensive physical tests further confirm strong transferability across scales, views, and scenarios. Attack demos are provided in the supplementary materials.

Min Dang, Gang Liu 0006, Jingqi Zhao, Adams Wai-Kin Kong, Nan Luo, Di Wang 0011

Infrared-visible image fusion for object detection (IVIF-OD) aims to utilize complementary information in the two modalities to synthesize new images with richer information to serve object detection. Most existing works focus on how to better fuse pixel-level details while ignoring object-related information required for detection and introducing redundant and object-irrelevant information in the fused images. To address the limitations of previous studies, this paper proposes a diffusion-based denoising fusion for object detection in infrared-visible images, termed DDFD. Specifically, DDFD treats image fusion as a diffusion-based denoising process to generate fused images that are informative yet non-redundant. Since visible imaging is easily affected by adverse conditions, DDFD exploits an image-adaptive enhancement (IAE) module that adaptively improves visible images to achieve better fusion. To extract key fusion features and remove redundancy, DDFD uses an image-aware noise estimator (INE) to determine the noise in the input infrared-visible images for promoting the diffusion denoising network. To take advantage of both the fusion network and object detection network, DDFD jointly optimizes them such that the fusion network can receive object information to improve the fused images, and the improved images can provide high-quality features to enhance object detection performance. Extensive experiments on the M3FD, DroneVehicle, and VEDAI public datasets reveal the superior object detection performance of DDFD and confirm the effectiveness of IVIF-based object detection under challenging weather conditions.

Yuhao Wang, Lingjuan Miao, Zhiqiang Zhou 0001, Lei Zhang 0243, Yajun Qiao

Infrared-visible image fusion (IVIF) has attracted much attention owing to the highly-complementary properties of the two image modalities. Due to the lack of ground-truth fused images, the fusion output of current deep-learning based methods heavily depends on the loss functions defined mathematically. As it is hard to well mathematically define the fused image without ground truth, the performance of existing fusion methods is limited. In this paper, we propose to use natural language to express the objective of IVIF, which can avoid the explicit mathematical modeling of fusion output in current losses, and make full use of the advantage of language expression to improve the fusion performance. For this purpose, we present a comprehensive language-expressed fusion objective, and encode relevant texts into the multi-modal embedding space using CLIP. A language-driven fusion model is then constructed in the embedding space, by establishing the relationship among the embedded vectors representing the fusion objective and input image modalities. Finally, a language-driven loss is derived to make the actual IVIF aligned with the embedded language-driven fusion model via supervised training. Experiments show that our method can obtain much better fusion results than existing techniques.

Mi Zheng, Guanglei Yang, Zitong Huang, Zhenhua Guo 0001, Kevin Han, Wangmeng Zuo

With the emergence of transformer-based architectures and large language models (LLMs), the accuracy of road scene perception has substantially advanced. Nonetheless, current road scene segmentation approaches are predominantly trained on closed-set data, resulting in insufficient detection capabilities for out-of-distribution (OOD) objects. To overcome this limitation, road anomaly detection methods have been proposed. However, existing methods primarily depend on image inpainting and OOD distribution detection techniques, facing two critical issues: (1) inadequate consideration of the objectiveness attributes of anomalous regions, causing incomplete segmentation when anomalous objects share similarities with known classes, and (2) insufficient attention to environmental constraints, leading to the detection of anomalies irrelevant to autonomous driving tasks. In this paper, we propose a novel framework termed Segmenting Objectiveness and Task-Awareness (SOTA) for autonomous driving scenes. Specifically, SOTA enhances the segmentation of objectiveness through a Semantic Fusion Block (SFB) and filters anomalies irrelevant to road navigation tasks using a Scene-understanding Guided Prompt-Context Adaptor (SG-PCA). Extensive empirical evaluations on multiple benchmark datasets, including Fishyscapes Lost and Found, Segment-Me-If-You-Can, and RoadAnomaly, demonstrate that the proposed SOTA consistently improves OOD detection performance across diverse detectors, achieving robust and accurate segmentation outcomes.

Xiangping Zheng, Xuan Feng, Bo Wu 0026, Bin Ren, Wei Li 0109, Xiuxin Hao, Xun Liang 0001, Bin Tang, Zhiwen Yu 0001

Cross-domain graph anomaly detection (GAD) aims to identify nodes that significantly deviate from normal patterns in unseen target domains, showing great potential in applications such as multimedia content security and financial risk control. However, existing methods often rely on semantic information trained on individual datasets, which makes it difficult to capture node commonalities across domains and limits generalization in complex multimedia environments. To address these challenges, we propose Zero-GAD, a universal Zero-shot Graph Anomaly Detection framework tailored for cross-domain scenarios. Zero-GAD leverages a novel de-semanticized strategy to train a unified detection model that can be directly applied to unseen domains without retraining or fine-tuning. The framework is built upon two key components: (1) a Global Information Unification Module, which projects graph data into the spectral domain and performs normalization to align the energy distribution in the frequency space; and (2) a Node-Neutralized Discrepancy Scoring Module that leverages the discrepancy between the original and reconstructed node representations to produce effective anomaly scores. Extensive experiments show that Zero-GAD achieves superior accuracy compared to existing models under a GAD setting.

Hongyu Jiang, Yuxin Huo, Sirou Sheng, Hong Tao, Chenping Hou

Traditional multi-view clustering methods rely on the cross-view sample alignment presumption to explore consistent and complementary information from multiple views. However, in real-world scenarios, sensor heterogeneity and decentralized data storage and processing frequently make this presumption violated, leading to the Unaligned Multi-view Clustering (UMC) problem. Although existing works have promoted the development of UMC, they have at least one of the following limitations, i.e., high computational complexity, inadequate use of high-order correlation and two-stage clustering. To address these limitations, we propose a Joint High-order Correlation Learning (JHCL) framework for scalable one-step unaligned multi-view clustering. Specifically, multi-order bipartite graphs are utilized to make fully use of intra-view high-order correlations. Then, based on a tensorial bipartite graph alignment and fusion model, inter-view high-order correlations are exploited simultaneously. In such manner, the learned consistent bipartite graph retains adequate structural information for accurate and fast clustering in one step. Extensive experiments on real-world datasets validate the superiority of JHCL in both clustering performance and computational efficiency. Code available: https://github.com/revolution6575/JHCL.git.

Ronghui Li, Lingxiao Han, Shi Shu, Yueyao Liu, Yukang Lin, Yue Ma 0016, Jie Guo, Ziwei Liu 0002, Xiu Li 0001

Existing LLM-based motion models fail to fully leverage large models' planning capabilities for motion-related tasks, exhibiting poor generalization, limited text-motion alignment, and an inability to perform multimodal condition joint driven motion generation. We argue that these issues arise from the modality gap and the highly coupled nature of motion tokens. To address this, we proposed the hybrid motion sentence, which is consistant of fine-grained motion decription and atomic body-part motion token that can bridge the gap between motion and text. To obtain a large corpus of hybrid motion sentences, we introduced a novel motion-to-text generation method that combines atomic motion operators with GPT-4o, resulting in 68.2 million fine-grained textual descriptions across diverse modalities. To reconstruct high-quality motion from hybrid sentences and make better motion-text alignment, we introduce Semantic-Aware Decoupled Motion Tokenization. Furthermore, we propose MotionUPG based on LLaMA, leveraging MotionWords dataset for both pretraining and instruction tuning. Our method achieves strong fine-grained text-motion alignment, impressive zero-shot motion generation, and is the first to support multimodal condition joint driven motion generation tasks.

Siyuan Zhang, Xiaoping Wang 0001, Jiang Li 0004, Weibin Feng, Xin Zhan, Hongzhi Huang

In robotics and autonomous driving, accurate depth estimation is vital yet challenging under dynamic scenes and extreme lighting. Conventional frame-based cameras offer rich context but suffer from motion blur and limited dynamic range, while event cameras provide high temporal resolution and dynamic range but lack global scene structure. Therefore, recent studies explore frame-event fusion depth estimation methods to leverage these two complementary modalities to achieve robust performance. However, due to the mismatch in temporal and spatial resolution, there is an inherent contradiction between high spatial resolution frames captured at sparse temporal intervals and event streams characterized by spatial sparsity but high temporal resolution, rendering cross-modal feature fusion ineffective. Moreover, the limited availability of frame-event depth datasets further undermines the model's generalization capability across different scenes. To address the above challenges, we propose HAFUNet, a Hierarchical Attention Fusion Network for depth estimation via frame-event fusion. Our method contains: (1) a pre-trained Dual-Stream Encoder (DSEer) to extract complementary features from frame and event inputs; (2) a Cross-modal Feature Interaction Module (CFIM) that aligns and fuses spatial-channel features across modalities; and (3) a Hierarchical Attention Decoder (HADer) that progressively refines depth predictions via attention-guided convolution. Experiments on synthetic and real-world datasets show that HAFUNet surpasses existing methods in depth accuracy and robustness. These results demonstrate the strength of our fusion strategy in diverse environments. Code is available at https://github.com/SiYZhangwh/HAFUNet.

Zhiwei Zhang 0005, Ruikai Xu, Weijian Zhang, Zhizhong Zhang 0001, Xin Tan 0002, Jingyu Gong, Yuan Xie 0006, Lizhuang Ma

In this paper, we present the first pinhole-fisheye framework for heterogeneous multi-view depth estimation, PFDepth. Our key insight is to exploit the complementary characteristics of pinhole and fisheye imagery (undistorted vs. distorted, small vs. large FOV, far vs. near field) for joint optimization. PFDepth employs a unified architecture capable of processing arbitrary combinations of pinhole and fisheye cameras with varied intrinsics and extrinsics. Within PFDepth, we first explicitly lift 2D features from each heterogeneous view into a canonical 3D volumetric space. Then, a core module termed Heterogeneous Spatial Fusion is designed to process and fuse distortion-aware volumetric features across overlapping and non-overlapping regions. Additionally, we subtly reformulate the conventional voxel fusion into a novel 3D Gaussian representation, in which learnable latent Gaussian spheres dynamically adapt to local image textures for finer 3D aggregation. Finally, fused volume features are rendered into multi-view depth maps. Through extensive experiments, we demonstrate that PFDepth sets a state-of-the-art performance on KITTI-360 and RealHet datasets over current mainstream depth networks. To the best of our knowledge, this is the first systematic study of heterogeneous pinhole-fisheye depth estimation, offering both technical novelty and valuable empirical insights.

Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Xinyi Wang

Multi-modal visual tracking leverages complementary sensor information to enhance robustness under challenging conditions. However, the security of multi-modal tracking systems remains largely unexplored. Existing attacks primarily target single-modal trackers or independently disrupt each modality, failing to exploit the inherent feature interactions and fusion mechanisms that define multi-modal tracking. As a result, these methods exhibit limited attack effectiveness and fail to assess multi-modal tracking systems' vulnerabilities accurately. Understanding these security risks is crucial, as adversarial threats could lead to severe failures in safety-critical applications. To address these challenges, a feature-aware adversarial attack, termed FA3T is proposed. It is designed to explicitly disrupt feature extraction and cross-modal alignment, thereby weakening the fusion process that multi-modal trackers rely on. To achieve this, a Frequency-Spatial Feature Separation (FSFS) module is constructed to perturb feature representations at multiple levels, weakening the modality-complementary advantages of multi-modal tracking. Furthermore, a Target Confusion Attack (TCA) module is devised to manipulate the target-background-template relationships, making it increasingly difficult for the tracker to distinguish the true target, significantly impairing tracking performance. Extensive experiments on five benchmark datasets (i.e., LasHeR, RGBT234, DepthTrack, VOT-RGBD2022, VisEvent) across three different modalities (RGB-T, RGB-D, and RGB-E) demonstrate that our attack substantially degrades state-of-the-art multi-modal trackers, exposing their susceptibility to adversarial threats.

Mianzimei Yang, Zhipeng Zhou, Jin Zhang 0035, Yuanhao Pu, Hong Xie 0004, Defu Lian

Deep long-tailed recognition (DLTR) has garnered increasing attention due to the inherent imbalance in many real-world problems (e.g., multimedia processing). Recently, some multi-objective optimization (MOO)-based solutions have been proposed to address conflicts during representation learning in DLTR. However, these methods face two primary challenges: (1) their effectiveness is subject to the power of MOO, which is arguable in recent literature, and (2) MOO approaches are resource-intensive due to frequent gradient operations. In this paper, we propose a novel approach: conflict-Buffering OptimizatiOn by Symmetry Teleportation (BOOST), which avoids altering complicated gradient combinations as previous methods did. A major challenge in this approach is the absence of off-the-shelf symmetry teleportation algorithms suitable for modern deep neural networks. To address this, we cast symmetry teleportation as the optimization of low-rank adaptation (LoRA). Specifically, we first divide categories into multiple groups and detect conflicts among them. When a conflict arises, we employ LoRA to identify an alternative point on the same loss level set, reducing conflicts and facilitating balanced optimization. To achieve this, we decouple symmetry teleportation into two objectives-loss invariance and balanced gradient maximization-and design corresponding objectives for LoRA optimization. Besides, we propose a trajectory reuse strategy to continually benefit from advanced optimizers. Extensive experiments demonstrate that BOOST achieves state-of-the-art performance across multiple mainstream DLTR datasets.

Chengzhou Li, Xiaokang Liu, Qi Jia 0001, Jinyuan Liu 0001, Zhiying Jiang, Longhan Feng, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001

Sonar image recognition is a key technology in underwater exploration systems. Compared with natural images, sonar images have fewer texture details and are easily affected by heavy noise, making it more challenging for specialists to distinguish the subtle differences among classes. In view of this, studying fine-grained classification methods for sonar images with scarce annotations is of significant importance. To address this issue, we propose a Physics-Guided Teacher-Student (PGTS) framework to explore the unique physical information of sonar images while simultaneously mitigating the effects of limited annotations. First, PGTS reconstructs sonar signals through physical simulation and a specially designed physics-guided feature generation module, which allows it to bypass the time-consuming physical simulation during inference. Then, we design a multi-modal teacher model combines the reconstructed sonar signals and sonar images to extract discriminative features to generate robust pseudo labels for fine-grained target categories. Finally, the knowledge is transferred to a single-modal student model through consistency loss. Under the joint constraints of the teacher model and the reconstructed sonar physical signals, the student model continuously improves its performance in annotation-scarce scenarios. Notably, when merely 1% of the data is labeled, our method outperforms other state-of-the-art approaches by 12.46% in terms of accuracy.

Xuanming Jiang, Baoyi An 0001, Zhengwei Zou, Dingyu Nie, Jialie Shen 0001, Xueming Qian, Guoshuai Zhao

In cutting-edge domains such as unmanned aerial vehicles and autonomous driving, edge-based audio-visual systems struggle to strike an optimal balance between complexity and performance. Unlike prevailing approaches that typically rely on pruning and knowledge distillation to streamline unimodal or hybrid models, we propose the Bio-Inspired Multimodal Network (BIMNet), which achieves an efficient audio-visual shared architecture. BIMNet integrates bio-inspired audio-visual modules that emulate the hierarchical sensory integration observed in nocturnal birds, to replicate equivalent biological information flow for both multiscale night vision and noise-adaptive hearing. Experimental findings show that BIMNet achieves superior performance and efficiency in diverse image datasets (varying in spatial scales and lighting conditions), audio datasets (encompassing various types of human and environmental sound), and audio-visual joint event detection tasks. Project support is available at: https://github.com/Mental-Scholar/BIMNet.

Disen Hu, Xun Jiang 0001, Zhe Sun 0009, Hao Yang 0015, Chong Peng, Peng Yan, Heng Tao Shen, Xing Xu 0001

Multimodal learning, which has been given great significance recently, may face the challenge of the imbalanced multimodal phenomenon, which leads to the insufficient optimization of both multimodal and unimodal objectives. The core problem lies in the optimization conflicts between the above optimization objectives, resulting in the diverse updating directions and strengths that cause antagonism between them. In this paper, we mathematically analyze the optimization processes of imbalanced multimodal learning in the hyperspaces from a novel geometric perspective. Additionally, based on our theoretical analysis, we defined the volumes of the gradients constructed parallel polyhedron in the hyperspace to quantify the misalignment between the optimization objectives. Subsequently, we proposed the Geometric Gradient Divergence Modulation (GGDM), which leverages the volumes of gradient polyhedron to perform gradient modulation, encouraging alignment among gradients and promoting a synergistic optimization effect. Lastly, we evaluate our GGDM on five widely used multimodal benchmarks, where RGB image, optical flow, text, image, video and audio are involved. Our method achieved state-of-the-art performance compared to other imbalanced multimodal learning methods. Our code is available at: https://github.com/ConstantineWayne/GGDM.