Vision Transformers (ViTs) show promising potential in various multimedia application scenarios. To facilitate their deployment on resource-constrained devices, token pruning and merging have been introduced. However, existing token compression methods focus solely on the abstract features exhibited by high-dimensional tokens after patch embedding, resulting in information loss when evaluating token importance. In this paper, we propose a novel Training-Free Adaptive Token Merging (TF-ATM) method by exploring the intrinsic properties of images themselves. Our TF-ATM is inspired by the observation that the characteristics of patches can intuitively reflect the redundancy level of images. Based on the observation, we develop a method that is mathematically formulated to merge tokens corresponding to patches that are close to the Median presentation in the Frequency domain (MF). The principle behind our merging is that patches close to MF can be replaced by their surrounding ones, and thus removing them does not impair performance. Besides, we experimentally show that patches farther from MF contain more important information, which can be leveraged to capture the object of interest accurately. Without any retraining, TF-ATM leads to significant improvements over the state-of-the-arts (SOTAs), with similar FLOPs (floating point operations). For example, we achieve a 44.5%-FLOPs reduction with only a small loss of 0.39% in top-1 accuracy for the MAE-H model on ImageNet dataset, superior to comparison approaches that require meticulous fine-tuning.
论文检索
输入标题、作者或关键词,从 1,620 篇学术成果中精准定位
Cross-Embodiment Learning (CEL) aims to train a generalist policy model by integrating large-scale compositional interactions of heterogeneous agents and environments. However, the inherent conflict between the unbounded space of agent-environment combinations and a single unified policy model hinders generalization to unseen combinations. To address this challenge, we propose a novel Mixture of Disentangled Prototypes (MoDP) method to improve the compositional generalization in CEL. The key idea is to introduce a finite prototype space that bridges the gap between unbounded agent-environment combinations and a single policy model. Specifically, we design a dual-headed autoencoder and a compositional reconstruction loss to disentangle agent and environment features from interaction data, and map them into respective prototype spaces. We then introduce a connection-sensitivity-based pruning method to extract sub-networks from the pre-trained policy model, forming policy prototypes associated with specific agent-environment prototype pairs. Finally, a parameter-free routing mechanism adaptively integrates relevant policy prototypes for each input composition. Experiments in both standard and compositional settings demonstrate the effectiveness of our MoDP in enhancing the generalization capability of pre-trained policies.
Low-light image enhancement aims to improve the visibility of degraded images to better align with human visual perception. While diffusion-based methods have shown promising performance due to their strong generative capabilities. However, their unidirectional modelling of degradation often struggles to capture the complexity of real-world degradation patterns, leading to structural inconsistencies and pixel misalignments. To address these challenges, we propose a bidirectional diffusion optimization mechanism that jointly models the degradation processes of both low-light and normal-light images, enabling more precise degradation parameter matching and enhancing generation quality. Specifically, we perform bidirectional diffusion-from low-to-normal light and from normal-to-low light during training and introduce an adaptive feature interaction block (AFI) to refine feature representation. By leveraging the complementarity between these two paths, our approach imposes an implicit symmetry constraint on illumination attenuation and noise distribution, facilitating consistent degradation learning and improving the model's ability to perceive illumination and detail degradation. Additionally, we design a reflection-aware correction module (RACM) to guide color restoration post-denoising and suppress overexposed regions, ensuring content consistency and generating high-quality images that align with human visual perception. Extensive experiments on multiple benchmark datasets demonstrate that our method outperforms state-of-the-art methods in both quantitative and qualitative evaluations while generalizing effectively to diverse degradation scenarios.Code
Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal Fusion
Vision-Language-Action (VLA) systems are crucial for autonomous decision-making in embodied intelligence. While current systems have advanced the instruction-following capabilities, their limited spatial perception often leads to suboptimal performance for mobile manipulation tasks in unstructured environments. To address this challenge, we propose Uni-Sight, an end-to-end VLA system for robust mobile manipulation. Uni-Sight unifies decision-making, perception, and control through joint training, enabling synchronized cross-component optimization. Within the system, we introduce Latent Feature Aligner (LFA) that ensures accurate target localization by aligning multi-view data. Specifically, we develop Domain Transfer Policy (DTP), a hierarchical policy constrained by LiDAR-guided spatial priors, which ensures 3D spatial understanding with limited visual coverage. Extensive experiments on 20 real-world mobile manipulation tasks demonstrate the high task success rate and robust execution performance of Uni-Sight. Our Uni-Sight achieves a 3.04× the success rate of existing methods, and exhibits superior generalization in both long-horizon and zero-shot scenes. Code and dataset are publicly available at https://github.com/trantor2nd/Uni-Sight.
Taming Anomalies with Down-Up Sampling Networks: Group Center Preserving Reconstruction for 3D Anomaly Detection
PDF ↗Reconstruction-based methods have demonstrated very promising results for 3D anomaly detection. However, these methods face great challenges in handling high-precision point clouds due to the large scale and complex structure. In this study, a Down-Up Sampling Networks (DUS-Net) is proposed to reconstruct high-precision point clouds for 3D anomaly detection by preserving the group center geometric structure. The DUS-Net first introduces a Noise Generation module to generate noisy patches, which facilitates the diversity of training data and strengthens the feature representation for reconstruction. Then, a Down-sampling Network (Down-Net) is developed to learn an anomaly-free center point cloud from patches with noise injection. Subsequently, an Up-sampling Network (Up-Net) is designed to reconstruct high-precision point clouds by fusing multi-scale up-sampling features. Our method leverages group centers for construction, enabling the preservation of geometric structure and providing a more precise point cloud. Extensive experiments demonstrate the effectiveness of our proposed method, achieving state-of-the-art (SOTA) performance, with an Object-level AUROC of 79.9% and 79.5% and a Point-level AUROC of 71.2% and 84.7% on the Real3D-AD and Anomaly-ShapeNet datasets, respectively.
Accessibility of multimedia content for all users, particularly blind and low-vision individuals (BLVIs), remains a significant challenge. While screen readers assist BLVIs by converting text to speech via Alt-Text and image descriptions, these methods are inherently text-based and struggle to convey spatial and graphical information effectively. To help this, we propose a framework that converts graphical components into tactile graphics rendered on a refreshable pin array. Our framework leverages on-device AI models to generate tactile representations without transmitting personal data. It thereby minimizes processing time and mitigates privacy concerns. The benchmark test showed that our on-device AI outperformed GPU servers (RTX 4090) operating in an intranet environment. To optimize the tactile output and evaluate the system's effectiveness on media accessibility, we conducted a series of user studies with three different use case scenarios. First, we derived the optimal threshold values for edge detection in tactile graphics, which resulted in 70 on a 0-255 scale. Then, we compared the proposed system to a vision language model (VLM; GPT-4o). The results indicated that our proposed framework is more effective regarding both information delivery and subjective satisfaction. The proposed framework can be directly applied to several visual media accessibility scenarios, with the benefits of using local AI, such as privacy protection, personalization, and cost-effectiveness.
As the use of dynamic point clouds (DPCs) expands in immersive media settings including augmented and virtual reality, it has become more important than ever to have precise and scalable methods for quality evaluation. However, most existing objective Point Cloud Quality Assessment (PCQA) methods focus on static content and fail to capture the temporal dynamics and multimodal perceptual cues inherent in dynamic scenarios. In this work, we propose a no-reference dynamic PCQA framework that integrates both geometric and visual modalities with global temporal modeling for perceptually aligned quality prediction. For the 3D modality, we extract localized spatio-temporal features using a time-aware point cloud encoder that incorporates the normalized frame index as an additional input channel. In parallel, we generate two complementary projections per frame and extract visual features using a pre-trained convolutional network. A dynamic gating network adaptively weights the contributions of the two modalities at each time step. These weighted features are fused and passed to a temporal transformer, which captures long-range temporal dependencies to regress the final quality score. Comprehensive tests on benchmark datasets reveal that our approach surpasses existing full-reference and no-reference PCQA techniques, demonstrating its efficacy in assessing the quality of dynamic point clouds.
Teleportation dominates Virtual Reality (VR) locomotion, known for its efficiency and ease of use, particularly in small physical spaces where users have limited room to move. However, is efficiency the only metric that matters? In this study, we challenge this well-established technique by comparing teleportation with an alternative approach for navigating constrained spaces: Walking-with-Portals. We conducted a user study (N = 24) comparing both techniques, collecting data on orientation, cybersickness, presence, and immersion. Our results reveal a distinct trade-off: Teleportation proved significantly faster, more efficient for pathfinding, and induced less cybersickness. Conversely, Walking-with-Portals significantly enhanced users' sense of spatial presence and perceived control over their movement. While a majority preferred Teleportation for the task's efficiency, 91.7% of participants identified Walking-with-Portals as the more immersive technique. Additionally, we establish key design guidelines for intuitive portal placement that allow a large virtual environment to fit within the small play areas (about 2.5 × 2.5 meters) common to home VR users. Our findings suggest that while teleportation remains the default solution in VR locomotion, Walking-with-Portals provides unique benefits that should not be overlooked. This work invites a reflection on locomotion in VR, arguing that the future of VR navigation may go beyond teleport.
Vietnamese street food has gained global popularity through platforms like YouTube, where creators are incentivized to post videos on specific topics, including street food. Efficient video editing has become essential for YouTubers. This study analyzed 1,507 videos and gathered insights from 213 respondents to identify the factors that shape viewer preferences and highlight key video features for creators. The study's key finding was the strong synergy between visual elements and linguistic richness, emphasizing the connection between storytelling and what appears on screen. Using machine learning, we examined visual, linguistic, and acoustic features, achieving a 70.5% accurate predictive model. Compared to the baseline 50%, these insights led to personalized recommendations for street food content creators, offering strategies to enhance viewer engagement. Our automated recommendation system bridges data-driven insights with content creation, elevating Vietnamese street food on the global stage and celebrating the blend of local culture and technology.
A high effort in Quality of Experience (QoE) research has been put into subjective assessment to determine the perceived quality of video. Most laboratory experiments follow guidelines from the ITU-T Recommendations, which suggest Absolute Category Rating (ACR) as a method to conduct such experiments. However, this method of video assessment radically limits confounding variables, is unnatural, and is far from how people cope with quality degradation and the cues they receive in these situations. This paper addresses this issue and proposes a more realistic subjective experiment based on the participant's behavior. Instead of passively rating the degraded quality of silent videos, we created the possibility to react to the annoying quality of chosen Netflix movies and reward participants by increasing the quality to the best possible. To cope with the data obtained, we adapted the method of fitting psychometric functions known in neuroscience, auditory science, animal science, and psychology. As a result, we obtained a more comprehensive image of the participants and their perceived quality, including their consistency, lapses, and differences. We estimated the parameters using Maximum Likelihood Estimation (MLE) and evaluated the goodness-of-fit of three S-shaped functions: Weibull, cumulative normal, and logistic. In the end, as the most common, we analyze the basic properties of the fitted functions, such as the Point of Subjective Equality (PSE), confidence intervals, and slope (β). We examine their role in describing individual differences among 34 subjects.
Inverse-Tone-Mapped High Dynamic Range Video Quality Assessment (ITM-HDR VQA) plays a pivotal role in evaluating the visual quality of ITM-enhanced HDR videos. The research community tackles this issue from the dataset and method perspectives. However, current ITM-HDR VQA datasets exhibit three key limitations: narrow scene diversity, partial HDR format representation, and inadequate distortion coverage; existing methods face challenges in HDR and SDR domain discrepancy and insufficient ITM-induced quality feature extraction. To bridge these gaps, we introduce a comprehensive ITM HDR Video Quality Assessment dataset tailored to Broadcast Television (BT-ITM-VQA), along with a novel SDR-Referenced Bidirectional Quality Interaction (SDR-R-BQI) method. The BT-ITM-VQA dataset features rich broadcast scenes, multiple HDR-format support of Hybrid Log-Gamma (HLG) and Perceptual Quantizer (PQ), and real-world distortions induced by super-resolution and deinterlacing, providing a systematic foundation for ITM-HDR VQA model development and validation. The SDR-R-BQI method effectively mitigates HDR and SDR discrepancies through luminance dynamic range alignment and color gamut alignment, and then extracts ITM-induced quality alterations by bidirectional, cross-quality-based computation in a unified feature space. Extensive validation on four datasets demonstrates the effectiveness of our newly constructed dataset and proposed method.
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
PDF ↗The furnishing of multi-modal large language models (MLLMs) has led to the emergence of numerous benchmark studies, particularly those evaluating their perception and understanding capabilities. Among these, understanding image-evoked emotions aims to enhance MLLMs' empathy, with significant applications such as human-machine interaction and advertising recommendations. However, current evaluations of this MLLM capability remain coarse-grained, and a systematic and comprehensive assessment is still lacking. To this end, we introduce EEmo-Bench, a novel benchmark dedicated to the analysis of the evoked emotions in images across diverse content categories. Our core contributions include: 1) Regarding the diversity of the evoked emotions, we adopt an emotion ranking strategy and employ the Valence-Arousal-Dominance (VAD) as emotional attributes for emotional assessment. In line with this methodology, 1,960 images are collected and manually annotated. 2) We design four tasks to evaluate MLLMs' ability to capture the evoked emotions by single images and their associated attributes: Perception, Ranking, Description, and Assessment. Additionally, image-pairwise analysis is introduced to investigate the model's proficiency in performing joint and comparative analysis. In total, we collect 6,773 question-answer pairs and perform a thorough assessment on 19 commonly-used MLLMs. The results indicate that while some proprietary and large-scale open-source MLLMs achieve promising overall performance, the analytical capabilities in certain evaluation dimensions remain suboptimal. Our EEmo-Bench paves the path for further research aimed at enhancing the comprehensive perceiving and understanding capabilities of MLLMs concerning image-evoked emotions, which is crucial for machine-centric emotion perception and understanding. Our code and benchmark datasets are available at https://github.com/workerred/EEmo-Bench.
Evaluating the visual quality of autostereoscopic 3D displays is crucial for quantifying their stereoscopic viewing experience and optimizing display performance. Existing quality evaluation methods primarily predict the visual quality of autostereoscopic 3D displays by indirectly learning display parameter information from image content. However, these methods fail to explicitly model the relationship between display parameters and visual quality, thereby limiting their prediction accuracy. To address this problem, a Multimodal Parameter Perception Network (MPPNet)-based visual quality assessment method is proposed in this paper, which treats display parameters as textual modalities to explicitly establish their relationship with visual quality. To effectively understand the semantic information of display parameter texts, a Contrastive Language-Image Pretraining (CLIP)-based adaptive text encoder is proposed to generate robust semantic representations by capturing both general and domain-specific semantic embeddings. In parallel, a hierarchical vision encoder is adopted to extract visual representations from display images, which simulates the human binocular perception by capturing multi-level visual features from the left and right views. To achieve comprehensive cross-modal interaction, a mamba-based cross-modal fusion module is proposed to fuse textual and visual representations of display parameters by capturing both shallow and deep correlations. Extensive experimental results demonstrate that the proposed MPPNet achieves state-of-the-art performance in evaluating the visual quality of autostereoscopic 3D displays.
Federated learning remains vulnerable to backdoor attacks through malicious parameter updates, with existing defenses limited by homogeneous data assumptions or reliance on gradient anomaly detection. We reveal that FedAvg's critical flaw lies in malicious feature extractor propagation: aggregating poisoned extractors degrades defense accuracy to <70% across five benchmarks, while benign extractors with poisoned headers retain an average of 89.36% defense accuracy. Therefore, we propose FeatShield, a feature-space isolation framework that prevents backdoor propagation via non-aggregated local extractors trained on clean client data. FeatShield introduces 1) variance-aware alignment, adaptively balancing client-specific features and global consistency using local variance metrics, and 2) adversarial feature synthesis, generating non-linear synthetic features via GAN to enhance the global prediction header's generalization on main tasks. Extensive experiments on eight real-world datasets show that FeatShield achieves the best defense performance. For instance, under heterogeneous data (Dirichlet β=0.5) and strong attacks (50% malicious clients), FeatShield achieves 99.26-99.89% defense accuracy and main task accuracy exceeding FedAvg by 1.32-5.70%, demonstrating its superior resistance to backdoor attacks without sacrificing the benign performance.
Three-dimensional (3D) light field displays (LFDs) provide immersive visual experiences and have attracted increasing attention. However, visual fatigue remains an important concern when users watch 3D LFDs which limits their development and application. In this paper, we propose a comprehensive methodology that integrates subjective and objective data to establish a robust dataset and employs eye movement data for systematically investigating visual fatigue in 3D LFDs. Firstly, a multimodal dataset is constructed by integrating subjective fatigue scores and objective eye movement data collection. Then, we propose the Deep Correlation Data Analysis Model (DCDAM), which uses Spearman's rank correlation coefficient to analyze correlations between key objective metrics and subjective fatigue curves, validating the effectiveness of these metrics. Furthermore, to comprehensively assess visual fatigue, we develop a specialized model, the Temporo-Spatial Synergy Network (TSSNet), which uses temporal and spatial eye movement features to predict subjective fatigue curves. Through validation across diverse videos, the model achieves R² > 0.98 (±0.005) and RMSE of 0.02 (±0.05) between actual and predicted values, demonstrating high precision and valid generalization across different video content. The proposed model provides a foundational framework for future research on visual fatigue assessment tasks of 3D LFDs.
Shadows are a common factor degrading image quality. Single-image shadow removal (SR), particularly under challenging indirect illumination, is hampered by non-uniform content degradation and inherent ambiguity. Consequently, traditional methods often fail to simultaneously recover intra-shadow details and maintain sharp boundaries, resulting in inconsistent restoration and blurring that negatively affect both downstream applications and the overall viewing experience. To overcome these limitations, we propose the DenseSR, approaching the problem from a dense prediction perspective to emphasize restoration quality. This framework uniquely synergizes two key strategies: (1) deep scene understanding guided by geometric-semantic priors to resolve ambiguity and implicitly localize shadows, and (2) high-fidelity restoration via a novel Dense Fusion Block (DFB) in the decoder. The DFB employs adaptive component processing-using an Adaptive Content Smoothing Module (ACSM) for consistent appearance and a Texture-Boundary Recuperation Module (TBRM) for fine textures and sharp boundaries-thereby directly tackling the inconsistent restoration and blurring issues. These purposefully processed components are effectively fused, yielding an optimized feature representation preserving both consistency and fidelity. Extensive experimental results demonstrate the merits of our approach over existing methods. Our code can be available on https://github.com/VanLinLin/DenseSR
ExplorAR: Assisting Older Adults to Learn Smartphone Apps through AR-powered Trial-and-Error with Interactive Guidance
PDF ↗Older adults tend to encounter challenges when learning to use new smartphone apps due to age-related cognitive and physical changes. Compared to traditional support methods such as video tutorials, trial-and-error allows older adults to learn to use smartphone apps by making and correcting mistakes. However, it remains unknown how trial-and-error should be designed to empower older adults to use smartphone apps and how well it would work for older adults. Informed by the guidelines derived from prior work, we designed and implemented ExplorAR, an AR-based trial-and-error system that offers real-time and situated visual guidance in the augmented space around the smartphone to empower older adults to explore and correct mistakes independently. We conducted a user study with 18 older adults to compare ExplorAR with traditional video tutorials and a simplified version of ExplorAR. Results show that the AR-supported trial-and-error method enhanced older adults' learning experience by fostering deeper cognitive engagement and improving confidence in exploring unknown operations.
Spatial audio playback defines immersive listening. However, objective evaluation methods for perceptual dimensions like sound field and sound image remain underdeveloped, hindered by the lack of fine-grained spatial audio datasets and the neglect of echoes and reverberation in diverse playback conditions. To address these challenges, we propose MESA, a multi-modal evaluation framework for spatial audio systems, and introduce PSA-MOS, a high-quality multi-scene spatial audio dataset. Specifically: 1) PSA-MOS provides 50 hours of high-quality spatial audio recordings spanning 6 playback scenarios and 7 device types, with detailed localization annotations and fine-grained MOS ratings across four perceptual dimensions. 2) We develop SAE-Encoder, a spatial audio encoder that captures both acoustic-spatial cues and fine-grained perceptual patterns. 3) MESA integrates visual scene context to enhance evaluation robustness through echo and reverberation modeling. Experimental results demonstrate that SAE-Encoder achieves superior performance in SELD tasks. With a two-stage training strategy, MESA exhibits strong correlation with human perceptual assessments, effectively guiding spatial audio quality optimization. The demos are available at https://david-pigeon.github.io/mesaDemo.
Large Language Model (LLM)-based agents have demonstrated strong capabilities across a wide range of tasks, and their application in the medical domain holds particular promise due to the demand for high generalizability and reliance on interdisciplinary knowledge. However, existing medical agent systems often rely on static, manually crafted workflows that lack the flexibility to accommodate diverse diagnostic requirements and adapt to emerging clinical scenarios. Motivated by the success of automated machine learning (AutoML), this paper introduces a novel framework for the automated design of medical agent architectures. Specifically, we define a hierarchical and expressive agent search space that enables dynamic workflow adaptation through structured modifications at the node, structural, and framework levels. Our framework conceptualizes medical agents as graph-based architectures composed of diverse, functional node types and supports iterative self-improvement guided by diagnostic feedback. Experimental results on skin disease diagnosis tasks demonstrate that the proposed method effectively evolves workflow structures and significantly enhances diagnostic accuracy over time. This work represents the first fully automated framework for medical agent architecture design and offers a scalable, adaptable foundation for deploying intelligent agents in real-world clinical environments.
Recent advancements in visual multi-task learning (MTL) have sparked significant interest. However, existing dense prediction MTL methods predominantly rely on single-modality image data, limiting their performance due to the absence of complementary knowledge from other modalities. Additionally, different dense tasks exhibit heterogeneous preferences during information decoding, posing a critical challenge in effectively allocating multi-scale encoded features. To address these limitations, we propose CLIP-MT, a Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction. Specifically, to enrich task-shared image features with multi-modal knowledge, we introduce a novel CLIP-Guided Global Feature Enhancer (CGGF), which leverages aligned text-image information to augment object-level representations through a dual-path feature fusion architecture. Furthermore, to tackle the task-specific scale preference problem, we design an Adaptive Scale Selection Gate (ASSG), a learnable gating mechanism that dynamically selects high- or low-scale features based on task-specific demands. Finally, we integrate multi-modal and multi-scale information through a Task-Aware Feature Fusion Module (TAFF). Extensive experiments on the NYUDv2 and PASCAL-Context datasets demonstrate that CLIP-MT achieves state-of-the-art performance, outperforming existing methods across multiple dense prediction tasks.