Vision-language models (VLMs) has demonstrated impressive cross-modal alignment. However, their internal mechanisms of associating text concepts with visual patterns remain opaque. This opacity raises a critical question: What visual patterns do VLMs inherently associate with text concepts? Current methods for decoding representations of VLMs often produce suboptimal outputs, hindering to probe the clear visual patterns. To address this, we introduce Generative Semantic Probing (GSP), a novel training-free framework that synthesizes images to probe the implicit semantic preferences of VLMs. Our method generates visual patterns that maximize the similarity to the target text embeddings, through three core components: (1) Hierarchical Feature Decomposition, which decomposes the image generation across multi-scale feature levels; (2) Feature Space Constraint, which constrains the optimization within semantically meaningful feature subspace; (3) Quality Assessment Module, which ensures the generation of visually plausible outputs. Experiments validate our method's strengths in high-fidelity image generation and interpretable model analysis. Beyond text-to-image generation, style transfer and image editing applications, our framework enables unprecedented visualization of VLMs' decision boundaries. By exposing implicit preferences and systematic biases in the cross-modal association, our work provides a valuable insight for both understanding and improvement of the vision-language alignment.
论文检索
输入标题、作者或关键词,从 7,537 篇学术成果中精准定位
Tumor spatial heterogeneity analysis requires precise correlation between Hematoxylin and Eosin (H&E) morphology and immunohistochemical (IHC) biomarker expression, yet current methods suffer from spatial misalignment in consecutive sections, severely compromising in situ pathological interpretation. In order to obtain a more accurate virtual staining pattern, We propose PRINTER, a weakly-supervised framework that integrates PRototype-drIven content and staiNing patTERn decoupling and deformation-aware adversarial learning strategies designed to accurately learn IHC staining patterns while preserving H&E staining details. Our approach introduces three key innovations: (1) A prototype-driven staining pattern transfer with explicit content-style decoupling; and (2) A cyclic registration-synthesis framework GapBridge that bridges H&E and IHC domains through deformable structural alignment, where registered features guide cross-modal style transfer while synthesized outputs iteratively refine the registration;(3) Deformation-Aware Adversarial Learning: We propose a training framework where a generator and deformation-aware registration network jointly adversarially optimize a style-focused discriminator. Extensive experiments demonstrate that PRINTER effectively achieves superior performance in preserving H&E staining details and virtual staining fidelity, outperforming state-of-the-art methods. Our work provides a robust and scalable solution for virtual staining, advancing the field of computational pathology.
Scalable Vector Graphics (SVG) has become an indispensable technology in front-end development and UI/UX design, due to its inherent advantages in scalability, editability, and rendering efficiency. In the creation of vector graphics, while expressing creative concepts is straightforward, translating them into precise digital artworks is often challenging and time-consuming. To overcome this technical bottleneck and achieve intelligent conversion from concept to final product, we have constructed SVG-1M, a large-scale dataset of high-quality SVG samples with paired textual descriptions. Through innovative data augmentation and annotation processes, we built precisely aligned ''Text instruction-SVG code'' training pairs, with a subset enhanced by Chain-of-Thought (CoT) annotations. This provides rich semantic supervision signals for model learning. Based on this dataset, we propose SVGen, an end-to-end generative model capable of directly converting natural language descriptions into SVG code. This design addresses the challenges of generating semantically accurate vector graphics while preserving complete structural information. We explored various training strategies and introduced a progressive curriculum learning approach, optimized with reinforcement learning algorithms. Notably, this study innovatively applies the CoT paradigm to vector graphics generation, effectively enhancing both the accuracy and interpretability of SVG synthesis. Experimental validation demonstrates that SVGen exhibits significant advantages over general large models in terms of SVG generation quality, while also surpassing optimization-based rendering methods in generation efficiency. The proposed method enables intelligent conversion between natural language and vector graphics, enabling novel workflows like real-time AI-assisted design iteration. Code, model, and data is released at: https://github.com/gitcat-404/SVGen
In low-light image enhancement, Retinex-based deep learning methods have garnered significant attention due to their exceptional interpretability. These methods decompose images into mutually independent illumination and reflectance components, allows each component to be enhanced separately. In fact, achieving perfect decomposition of illumination and reflectance components proves to be quite challenging, with some residuals still existing after decomposition. In this paper, we formally name these residuals as inter-component residuals (ICR), which has been largely underestimated by previous methods. In our investigation, ICR not only affects the accuracy of the decomposition but also causes enhanced components to deviate from the ideal outcome, ultimately reducing the final synthesized image quality. To address this issue, we propose a novel Inter-correction Retinex model (IRetinex) to alleviate ICR during the decomposition and enhancement stage. In the decomposition stage, we leverage inter-component residual reduction module to reduce the feature similarity between illumination and reflectance components. In the enhancement stage, we utilize the feature similarity between the two components to detect and mitigate the impact of ICR within each enhancement unit. Extensive experiments on three low-light benchmark datasets demonstrated that by reducing ICR, our method outperforms state-of-the-art approaches both qualitatively and quantitatively. Our code is available at: https://github.com/caoluyang0830/IRetinex.git.
Recent advances in text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality visual content with style and feature controlled. A fundamental challenge remains in simultaneously maintaining three critical properties of generated image sequences: (1) fine-grained style control, (2) strict image-prompt alignment, and (3) cross-image content coherence. To overcome the challenge, we leverage AnyStyleDiffusion to overcome the challenge. Specifically, we interpret any artistic style required by users on generated image as a feature in models' weight space. Interpolation between weight space obtains models expressing middle styles with linear transition. Hyper-receptive Motion Layers is proposed to align outputs of diverse weight spaces, operating as adaptive style modulators. These HRMLs are separated from interpolated diffusion models, leveraging zero-shot compatibility with existing model checkpoints. By employing Homogeneous Stable Diffusion, direct interpolation on weight space is avoided to improve synthesis efficiency. Comprehensive evaluations across personalized models demonstrate our method's superiority in generating content-coherent sequences with dynamic style transformations. Code will be released at https://github.com/shermandozer/AnyStyleDiffusion.git.
As short videos become a dominant medium for news dissemination, fake news videos pose increasing threats to public trust and information integrity. Existing methods primarily focus on learning multimodal representations to predict binary veracity labels, yet they overlook the use of external evidence, which is important for identifying more sophisticated fake news that subtly exploits psychological cues and cognitive biases. Moreover, these approaches do not provide fine-grained attribution labels, which are essential for interpretable misinformation governance. To address these limitations, we introduce EvidSV, the first comprehensive benchmark supporting evidence- and attribution-aware fake news video detection. Drawing inspiration from the human cognitive process of interpreting news-related content, we propose MUKE, a multi-view knowledge progressive enhancement learning framework. By jointly analyzing both the news content and supporting evidence, MUKE (1) facilitates the understanding of news semantics to (2) progressively refine shared domain knowledge, and (3) adaptively summarizes multi-view knowledge to assess news veracity. Extensive experiments demonstrate that MUKE consistently outperforms existing methods in both fake news detection and attribution, and generalizes effectively to previously unseen domains. Our code is available at https://github.com/zzeng1998/EvidSV.
With the rapid evolution of multimedia technologies and its widespread integration into education, adaptive multimedia learning has gained significant prominence. Cognitive diagnosis (CD) is pivotal in this domain, as it models students' cognitive states using practice data captured by multimedia learning applications. However, existing methods often simplify these states to mere proficiency on knowledge concepts. Constructivism in education emphasizes learning as a continuous cognitive development process, during which students' cognitive states become increasingly complex, involving not only their construction of concepts but also their construction of relations between concepts that have long been overlooked. To this end, we propose the Hierarchical Disentanglement of Cognitive States for Enhanced Cognitive Diagnosis (HDCD). Inspired by the Structure of Observed Learning Outcomes (SOLO) taxonomy, which categorizes cognitive development into core hierarchical levels (Multistructural, Relational, Extended Abstract), we introduce a hierarchical disentanglement strategy to define cognitive states aligned with each SOLO level: Intra-Concept Cognitive States, Relational Cognitive States, and Extended Cognitive States. Specifically, (i) At the multistructural level, intra-concept cognitive states are sampled from student's personalized cognitive distribution, representing the construction of individual concepts. (ii) At the relational level, inter-concept cognitive states are first sampled to represent the construction of relations between concepts. We then employ a hypergraph transformation to collaboratively update both intra-concept and inter-concept cognitive states, forming relational cognitive states. Considering that students' self-constructed knowledge systems involve multiple types of inter-concept relations, relational cognitive states are implemented under both undirected and directed relation views in this work, and then fed into local diagnostic functions, respectively. (iii) At the extended abstract level, outputs from the local diagnostic functions are fused using multi-view attention mechanisms, resulting in extended cognitive states, which integrate information from multiple relational views, are then fed into a global diagnostic function for final prediction. Extensive experiments on real-world datasets demonstrate the superior performance and interpretability of our HDCD.
We study how large vision models (LVMs) can predict food nutrition through lightweight and interpretable adapters---the machine learning modules the predictions of which could be understood by humans. We introduce novel nutrition adapters that use features extracted by pre-trained LVMs and output the so-called nutrition maps. Nutrition maps indicate the concentration of nutrition values per each image location. We use such an interpretable representation to obtain the nutrition targets as a sum of all nutrition concentrations on the maps. To understand our approach's generalization capability, we systematically analyze the behavior of our novel interpretable adapters leveraging different LVMs with different food image-nutrition datasets. Our lightweight approach delivers better or on-par performance than the state-of-the-art models on the Nutrition5k and the Nutritionverse-Real benchmarks. The code is provided at https://github.com/vitaly-emelianov/nutrition-adapters.
Effectively communicating uncertainty in ensemble hurricane forecasts poses a significant multimedia challenge, requiring the integration of spatial, temporal, and perceptual dimensions. We introduce SUVIS, a stereoscopic visualization system that encodes forecast ensembles into an immersive, layered media experience. SUVIS transforms multidimensional ensemble data into animated stereoscopic representations, mapping time to vertical depth, intensity to texture color, and forward speed to motion flow, while semi-transparent glyphs represent evolving impact areas. A progressive sampling strategy ensures spatial clarity across depth layers. Rendered on a glasses-free stereoscopic display, SUVIS frames uncertainty visualization as a media encoding problem, synthesizing motion, depth, and spatial abstraction to align with human perception. A user study with 51 participants demonstrates that SUVIS supports high accuracy in spatial tasks and enables interpretation of dynamic storm attributes. These results highlight the system's potential to advance perceptual uncertainty communication through multimedia representation and immersive visual encoding.
Facial image steganography is crucial for privacy-preserving media transmission. Traditional embedding methods degrade image quality and are vulnerable to steganalysis, while GAN-based non-embedding approaches lack controllability and realism. Diffusion-based methods using textual prompts face two key issues: (1) security risks from interpretable prompts and (2) poor preservation of facial details. This paper presents Featurized Denoising Diffusion Implicit Models (F-DDIM), a novel non-embedding steganography framework. First, F-DDIM replaces explicit textual prompts with implicit image-based encoding, enhancing security. Second, it selectively refines facial regions for natural and high-quality recovery through iterative reconstruction. Third, it enables indistinguishable encryption without secret key sharing via a novel sub-code embedding algorithm. Fourth, a refinement step post-decoding improves the clarity and accuracy of recovered facial image details. Experimental results demonstrate that F-DDIM achieves superior image fidelity and robustness against transmission interference.
Visual selective attention, driven by individual preferences, regulates human prioritization of visual stimuli by bridging subjective cognitive mechanisms with objective visual elements, thereby steering the semantic interpretation and hierarchical processing of dynamic visual scenes. However, existing models and datasets predominantly neglect the influence of subjective cognitive diversity on fixation behavior. Conventional saliency prediction models, typically employing segmentation approaches, rely on low-resolution imagery to generate saliency heatmaps, subsequently upscaled to native resolutions, which limiting their capacity to capture personalized attention patterns. Furthermore, MLLMs are constrained by factors such as hallucinations, making it very costly to strictly adhere to the expected format in tasks involving multiple point predictions, and achieving precise point positioning is challenging. To address these limitations, we present Subjective Personalized Attention for Ad vertisement Videos, namely SPA-ADV, a large-scale multimodal dataset capturing gaze behaviors from over 4,500 participants varying in age and gender with 486 videos. Furthermore, we propose PRE-MAP, a novel eye-tracking saliency model that characterizes Personalized visual disparities through Reinforcement learning-optimized Eye-tracking, built upon MLLMs and guided by Multi-Attribute user profiles to predict Points. To ensure MLLMs produce prediction points that are both format-correct and spatially accurate, we introduce Consistency Group Relative Policy Optimization (C-GRPO), inspired by the variability in eye movement points and Multi-Attribute profiles. Extensive experiments on SPA-ADV and other benchmarks demonstrate the effectiveness of our approach. The code and dataset are available at https://github.com/mininglamp-MLLM/PRE-MAP.
The proliferation of satellite-related commonsense data on the internet. Traditional analytical methods are challenging to integrate and effectively uncover implicit knowledge within them. However, current deep learning and LLM-based approaches often struggle with errors and hallucinations when performing multi-hop reasoning in domain-specific contexts. To address these limitations, we propose a novel multi-hop reasoning framework for implicit commonsense mining. This framework aims to uncover the underlying meta-knowledge behind reasoning problems, thereby providing enhanced interpretability of the reasoning process. Specifically, we design an extraction-retrieval-principle multi-step reasoning method that generates different levels of meta-knowledge in stages to support the reasoning process effectively. We further design the mixture of expert knowledge graph construction to construct a satellite knowledge graph that supports multi-hop reasoning. Experimental results demonstrate that our approach outperforms baselines on satellite knowledge graph reasoning.
Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by leveraging visionlanguage models (VLMs). However, these methods often fall short in aligning free-form language commands with specific scene instances, due to limitations in both instance-level semantic consistency and instruction interpretation. We present OpenMap, a zero-shot open-vocabulary visual-language map designed for accurate instruction grounding in navigation tasks. To address semantic inconsistencies across views, we introduce a Structural-Semantic Consensus constraint that jointly considers global geometric structure and vision-language similarity to guide robust 3D instancelevel aggregation. To improve instruction interpretation, we propose an LLM-assisted Instruction-to-Instance Grounding module that enables fine-grained instance selection by incorporating spatial context and expressive target descriptions. We evaluate OpenMap on ScanNet200 and Matterport3D, covering both semantic mapping and instruction-to-target retrieval tasks. Experimental results show that OpenMap outperforms state-of-the-art baselines in zero-shot settings, demonstrating the effectiveness of our method in bridging free-form language and 3D perception for embodied navigation.
Aesthetic Image Cropping (AIC) aims to improve the visual appeal of images by removing redundant content while preserving attractive elements. Despite the encouraging progresses achieved in data-driven approaches, most existing models struggle to understand user intentions, particularly for diversified scenes with multiple subjects. Moreover, they can only provide cropping results without explanations, which further restricts their usability in real-world applications. Motivated by the above facts, we introduce InstructCrop : a multimodal large language model (MLLM)-based AIC framework, which can understand user instructions and provide explanatory reasons for cropping results. Specifically, we first build a multimodal Image Cropping Instruction Tuning (ICIT) dataset through a cost-effective paradigm by generating high-quality instruction tuning data based on the existing cropping datasets. Then, we embed dynamic domain knowledge into the cropping model by integrating cropping-aware experts of aesthetic assessment and composition classification. Finally, we adapt MLLMs to generate the cropping results and corresponding explanations. Quantitative and qualitative experiments on three benchmark datasets demonstrate that InstructCrop enables effective and interpretable image cropping, which aligns better with user intentions. Data and code are available at https://github.com/sxfly99/InstructCrop.
Offering diverse perspectives on a museum artifact can deepen visitors' understanding and help avoid the cognitive limitations of a single narrative, ultimately enhancing their overall experience. Physical museums promote diversity through visitor interactions. However, it remains a challenge to present multiple voices appropriately while attracting and sustaining a visitor's attention in the virtual museum. Inspired by recent studies that show the effectiveness of LLM-powered multi-agents in presenting different opinions about an event, we propose SimViews, an interactive multi-agent system that simulates visitor-to-visitor conversational patterns to promote the presentation of diverse perspectives. The system employs LLM-powered multi-agents that simulate virtual visitors with different professional identities, providing diverse interpretations of artifacts. Additionally, we constructed 4 conversational patterns between users and agents to simulate visitor interactions. We conducted a within-subject study with 20 participants, comparing SimViews to a traditional single-agent condition. Our results show that SimViews effectively facilitates the presentation of diverse perspectives through conversations, enhancing participants' understanding of viewpoints and engagement within the virtual museum.
Chinese calligraphy offers fruitful visual structure not found in abstract/ figurative paintings or photographic images. It makes it well-suited for studying how personality shapes aesthetic preference, yet few works have explored this link. This paper introduces the first computational framework that models the link between viewer personality and Kai Shu calligraphic preference. It collects a dataset of Kai Shu calligraphy images, user preference scores, and Big Five personality traits. It extracts 160 structural feature descriptors from eight categories, such as stroke curvature, layout, and whitespace. Regression and attribution methods reveal five patterns, such as visual structure predicts perceived style, and that traits like Openness and Neuroticism influence preference patterns. High-Openness users prefer balanced, clean layouts. Low-Neuroticism users favor lighter, irregular forms. Some personality-feature pairs follow inverted-U trends, where moderate structural complexity leads to higher preference. These results connect cognitive traits with visual structure and support interpretable, personality aware modeling of aesthetic response. Our findings support personalized style discovery and open up new directions for interest-driven aesthetic education and digital preservation. Code available at https://github.com/tianchengliu18/kai2trait.
Understanding cultural heritage through technology faces challenges in connecting with diverse audiences, especially when interpreting art across cultures. In this work, we present CultiVerse, a visual analytics system that leverages Large Language Models (LLMs) to support cross-cultural appreciation of Traditional Chinese Paintings (TCPs). CultiVerse operates within a mixed-initiative framework and guides users through three stages: extracting cultural context, aligning cross-cultural symbols, and extrapolating meaning in the viewer's cultural frame. By combining an interactive interface with LLM-powered analysis, the system enables deeper engagement with symbolic meanings and encourages serendipitous cross-cultural discoveries. Our approach bridges AI interpretation and human insight to foster mutual understanding in a multicultural setting. A curated TCP dataset supports exploration, while empirical evaluations confirm that CultiVerse enhances user understanding, interpretation accuracy, and cultural empathy.
Visual art understanding requires joint modeling of multiple perspectives and contextual inference rooted in cultural, historical, and stylistic knowledge. Recent multimodal large language models (MLLMs) demonstrate strong performance in generic captioning, primarily based on object recognition and training on large-scale generic data. They struggle in providing captions incorporating the multiple perspectives that fine art demands. In this work, we introduce ArtRAG, a novel training-free framework that integrates structured knowledge into a retrieval-augmented generation (RAG) pipeline for multi-perspective artwork explanation. ArtRAG automatically constructs an Art Context Knowledge Graph (ACKG) from domain-specific textual sources, organizing entities such as artists, themes, movements, and historical events into a rich, interpretable knowledge graph. At inference time, a multi-granular structured context retriever selects semantically and topologically relevant subgraphs to guide explanation generation. This approach enables MLLMs to produce contextually grounded, multi-perspective descriptions. Experiments on the SemArt and Artpedia datasets demonstrate that ArtRAG outperforms existing heavily trained baselines. Human evaluations further confirm ArtRAG's ability to generate coherent, informative, and culturally enriched interpretations of artworks.
Tombstones are historically and culturally rich artifacts, encapsulating individual lives, community memory, historical narratives and artistic expression. Yet, many tombstones today face significant preservation challenges, including physical erosion, vandalism, environmental degradation, and political shifts. In this paper, we introduce a novel multi-modal framework for tombstone digitization, aiming to improve the interpretation, organization and retrieval of tombstone content. Our approach leverages vision-language models (VLMs) to translate tombstone images into structured Tombstone Meaning Representations (TMRs), capturing both image and text information. To further enrich semantic parsing, we incorporate retrieval-augmented generation (RAG) to integrate externally dependent elements such as toponyms, occupation codes, and ontological concepts. Compared to traditional OCR-based pipelines, our method improves parsing accuracy from an F1 score of 36.1 to 89.5. Furthermore, we evaluate the model's robustness across diverse linguistic and cultural inscriptions, and simulate physical degradation through image fusion to assess performance under noisy or damaged conditions. Our work represents the first attempt to formalize tombstone understanding using large vision-language models, presenting implications for heritage preservation. The code and supplementary materials are available at: https://github.com/LastDance500/Tombstone-Parsing.
Recent advances in generative modeling have enabled the synthesis of high-quality artistic images. Nevertheless, systematic evaluation of generative models from an aesthetic standpoint is still lacking, which hinders progress in artistic image synthesis. Existing evaluation metrics, such as Fréchet Inception Distance (FID) and CMMD, struggle with aesthetic assessment: they rely on pretrained visual features that overlook nuanced artistic attributes and employ distance functions ill-suited for modeling the diverse, multi-modal distribution of artistic styles. To address these limitations, we propose ArtFRD, a metric specifically designed for generative aesthetic evaluation. Grounded in aesthetic theory, ArtFRD extracts visual features along four key aesthetic dimensions-brushstroke, composition, lighting, and color-to capture fine-grained artistic properties. To model the multi-modal nature of artistic styles, we adopt a Gaussian Mixture Model assumption and derive an efficient approximation of the Fisher-Rao distance, which serves as the final evaluation score. Extensive experiments demonstrate that ArtFRD aligns significantly better with human aesthetic judgments than existing metrics, even across a wide range of artistic styles. These results highlight its potential as a robust and interpretable foundation for future research in generative aesthetic evaluation.