论文检索

输入标题、作者或关键词,从 12,319 篇学术成果中精准定位

会议来源 全部会议

机器学习与综合 AI

自然语言处理

计算机视觉

数据挖掘与 Web

多媒体与图形学

未选择时检索全部会议
支持跨会议组合检索,PDF 均跳转至官方来源
12,319篇论文匹配“Datasets and Benchmarks”
第 191 / 616 页

Negin Ghamsarian, Raphael Sznitman, Klaus Schoeffmann, Jens Kowal

To meet the growing demand for systematic surgical training, wet-lab environments have become indispensable platforms for hands-on practice in ophthalmology. Yet, traditional wet-lab training depends heavily on manual performance evaluations, which are labor-intensive, time-consuming, and often subject to variability. Recent advances in computer vision offer promising avenues for automated skill assessment, enhancing both the efficiency and objectivity of surgical education. Despite notable progress in ophthalmic surgical datasets, existing resources predominantly focus on real surgeries or isolated tasks, falling short of supporting comprehensive skill evaluation in controlled wet-lab settings. To address these limitations, we introduce WetCat, the first dataset of wet-lab cataract surgery videos specifically curated for automated skill assessment. WetCat comprises high-resolution recordings of surgeries performed by trainees on artificial eyes, featuring comprehensive phase annotations and semantic segmentations of key anatomical structures. These annotations are meticulously designed to facilitate skill assessment during the critical capsulorhexis and phacoemulsification phases, adhering to standardized surgical skill assessment frameworks. By focusing on these essential phases, WetCat enables the development of interpretable, AI-driven evaluation tools aligned with established clinical metrics. This dataset lays a strong foundation for advancing objective, scalable surgical education and sets a new benchmark for automated workflow analysis and skill assessment in ophthalmology training. The dataset and annotations are publicly available in Synapse (https://www.synapse.org/Synapse:syn66401174/files/).

Guangxun Zhu, Shiyu Fan, Hang Dai, Edmond S. L. Ho

Large-scale high-quality 3D motion datasets with multi-person interactions are crucial for data-driven models in autonomous driving to achieve fine-grained pedestrian interaction understanding in dynamic urban environments. However, existing datasets mostly rely on estimating 3D poses from monocular RGB video frames, which suffer from occlusion and lack of temporal continuity, thus resulting in unrealistic and low-quality human motion. In this paper, we introduce Waymo-3DSkelMo, the first large-scale dataset providing high-quality, temporally coherent 3D skeletal motions with explicit interaction semantics, derived from the Waymo Perception dataset. Our key insight is to utilize 3D human body shape and motion priors to enhance the quality of the 3D pose sequences extracted from the raw LiDRA point clouds. The dataset covers over 14,000 seconds across more than 800 real driving scenarios, including rich interactions among an average of 27 agents per scene (with up to 250 agents in the largest scene). Furthermore, we establish 3D pose forecasting benchmarks under varying pedestrian densities, and the results demonstrate its value as a foundational resource for future research on fine-grained human behavior understanding in complex urban environments. The dataset and code will be available at https://github.com/GuangxunZhu/Waymo-3DSkelMo

Yuchen Zhang, Tailin Chen, Jiangbei Yue, Yueming Sun, Rahul Singh, Jianbo Jiao, Zeyu Fu

Hate speech poses a persistent threat to society, causing profound harm to both individuals and communities. Detecting such content is essential for promoting safer and more inclusive environments. While previous research has primarily focused on text-based or image-based hate speech detection, video-based hate detection remains relatively underexplored. A key barrier is the limited availability of high-quality video datasets. Existing hateful video datasets are typically limited in scale, diversity, and annotation depth, often labeling hateful content without further distinguishing between explicit and implicit forms. In this work, we present DeHate, which, to the best of our knowledge, is the largest hateful video dataset to date. DeHate comprises 6689 videos collected from two platforms and spanning six social groups. Each video is annotated with fine-grained labels that differentiate explicit, implicit, and non-hateful content, along with segment-level localization of hate, identification of contributing modalities, and specification of the targeted groups. Through detailed analysis of annotated videos across platforms, we reveal distinct patterns in how hateful content is conveyed, offering a comprehensive comparison between explicit and implicit hate in terms of their prevalence and characteristics. Furthermore, we benchmark state-of-the-art models, including both uni-modal and multi-modal architectures, and identify persistent challenges in detecting subtle and context-dependent forms of hate. Our findings highlight the importance of holistic and fine-grained hateful video datasets for advancing research in hate speech detection. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.

Gabriele Magrini, Niccolò Marini, Federico Becattini, Lorenzo Berlincioni, Niccolò Biondi, Pietro Pala, Alberto Del Bimbo

Small, fast, and lightweight drones present significant challenges for traditional RGB cameras due to their limitations in capturing fast-moving objects, especially under challenging lighting conditions. Event cameras offer an ideal solution, providing high temporal definition and dynamic range, yet existing benchmarks often lack fine temporal resolution or drone-specific motion patterns, hindering progress in these areas. This paper introduces the Florence RGB-Event Drone dataset (FRED), a novel multimodal dataset specifically designed for drone detection, tracking, and trajectory forecasting, combining RGB video and event streams. FRED features more than 7 hours of densely annotated drone trajectories, using 5 different drone models and including challenging scenarios such as rain and adverse lighting conditions. We provide detailed evaluation protocols and standard metrics for each task, facilitating reproducible benchmarking. The authors hope FRED will advance research in high-speed drone perception and multimodal spatiotemporal understanding.

Jinke Li, Jiarui Yu, Chenxing Wei, Hande Dong, Qiang Lin, Liangjing Yang, Zhicai Wang, Yanbin Hao

Unlike bitmap images, scalable vector graphics (SVG) maintain quality when scaled, frequently employed in computer vision and artistic design in the representation of SVG code. In this era of proliferating AI-powered systems, enabling AI to understand and generate SVG has become increasingly urgent. However, AI-driven SVG understanding and generation (U&G) remain significant challenges. SVG code, equivalent to a set of curves and lines controlled by floating-point parameters, demands high precision in SVG U&G. Besides, SVG generation operates under diverse conditional constraints, including textual prompts and visual references, which requires powerful multi-modal processing for condition-to-SVG transformation. Recently, the rapid growth of Multi-modal Large Language Models (MLLMs) have demonstrated capabilities to process multi-modal inputs and generate complex vector controlling parameters, suggesting the potential to address SVG U&G tasks within a unified model. To unlock MLLM's capabilities in the SVG area, we propose an SVG-centric dataset called UniSVG, comprising 525k data items, tailored for MLLM training and evaluation. To our best knowledge, it is the first comprehensive dataset designed for unified SVG generation (from textual prompts and images) and SVG understanding (color, category, usage, etc.). As expected, learning on the proposed dataset boosts open-source MLLMs' performance on various SVG U&G tasks, surpassing SOTA close-source MLLMs like GPT-4V. We release dataset, benchmark, weights, codes and experiment details on https://ryanlijinke.github.io/.

Shifu Xiong, Hang Chen 0001, Shi Cheng 0001, Kai Shen, Hengshun Zhou, Genshun Wan, Chenyue Zhang, Kewei Li, Jun Du 0002, Lirong Dai 0001

We present MISP-QEKS, the first large-scale tri-modal benchmark for query-by-example keyword spotting (QEKS). Specifically, MISP-QEKS comprises 610,000 enrollment-query pairs with real-world noise, and covers 9,830 keywords. We also introduce a cross-modal enrolment-query matcher (XEQ-Matcher) as the baseline, which computes wake-word likelihoods by pairwise similarity between pre-trained enrolment and query embeddings. We also propose two plug-and-play modules: a visual gating module (VGM) that filters noise using lip movements, and a multimodal alignment module (MAM) that enforces phone-level consistency across modalities. Experiments show that XEQ-Matcher delivers peak performance with visual enrollment and query, achieving 82.82% AUC/24.23% EER and 79.79% AUC/26.20% EER on in-vocabulary and out-of-vocabulary splits. Incorporating the VGM further improves performance by +1.24% AUC/-0.89% EER and +2.65% AUC/-1.43% EER, respectively. And adding the MAM on top of VGM yields an additional +1.88% AUC/-1.74% EER and +3.00% AUC/-3.28% EER, respectively. These results confirm that tri-modal fusion significantly enhances robustness and generalisation of QEKS. MISP-QEKS with code is available at https://github.com/coalboss/MISP-QEKS.

Sahar Nasirihaghighi, Negin Ghamsarian, Leonie Peschek, Matteo Munari, Heinrich Husslein, Raphael Sznitman, Klaus Schoeffmann

Recent advances in deep learning have transformed computer-assisted intervention and surgical video analysis, driving improvements not only in surgical training, intraoperative decision support, and patient outcomes, but also in postoperative documentation and surgical discovery. Central to these developments is the availability of large, high-quality annotated datasets. In gynecologic laparoscopy, surgical scene understanding and action recognition are fundamental for building intelligent systems that assist surgeons during operations and provide deeper analysis after surgery. However, existing datasets are often limited by small scale, narrow task focus, or insufficiently detailed annotations, limiting their utility for comprehensive, end-to-end workflow analysis. To address these limitations, we introduce GynSurg, the largest and most diverse multi-task dataset for gynecologic laparoscopic surgery to date. GynSurg provides rich annotations across multiple tasks, supporting applications in action recognition, semantic segmentation, surgical documentation, and discovery of novel procedural insights. We demonstrate the dataset's quality and versatility by benchmarking state-of-the-art models under a standardized training protocol. To accelerate progress in the field, we publicly release the GynSurg dataset and its annotations (https://ftp.itec.aau.at/datasets/GynSurge/).

Hieu Nguyen, Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le

We introduce OpenEvents V1, a large-scale benchmark dataset designed to advance event-centric vision-language understanding. Unlike conventional image captioning and retrieval datasets that focus on surface-level descriptions, OpenEvents V1 dataset emphasizes contextual and temporal grounding through three primary tasks: (1) generating rich, event-aware image captions, (2) retrieving event-relevant news articles from image queries, and (3) retrieving event-relevant images from narrative-style textual queries. The dataset comprises over 200,000 news articles and 400,000 associated images sourced from CNN and The Guardian, spanning diverse domains and time periods. We provide extensive baseline results and standardized evaluation protocols for all tasks. OpenEvents V1 establishes a robust foundation for developing multimodal AI systems capable of deep reasoning over complex real-world events. The dataset is publicly available at https://ltnghia.github.io/eventa/openevents-v1.

Guillaume Gautier, Xuemei Zhou, Thong Nguyen, Jack Jansen 0001, Louis Fréneau, Marko Viitanen, Uyen Phan, Jani Käpylä, Irene Viola 0001, Alexandre Mercat 等

Volumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP).

Ines Riahi, Abduljalil Radman, Zixin Guo, Rachid Hedjam, Jorma Laaksonen

Despite growing interest in Audio-Visual Question Answering (AVQA), existing datasets often suffer from limited diversity, rigid formats, and insufficient integration of audio and visual modalities. To address these limitations, we introduce Valor32k-AVQA v2.0, a large-scale dataset containing 28,863 real-world videos and over 225,000 QA pairs, designed to support diverse and realistic multimodal understanding. The dataset features both open-ended and multiple-choice questions, each annotated with the required modality ( visual, audio, or audio-visual ) and question category ( description, action, count, temporal, location, or relative position ). All annotations-including questions, answers, and metadata-are generated through a fully automated prompting pipeline using GPT-4o, with human validation performed on a representative sample to ensure quality. We benchmark a few state-of-the-art models, with additional evaluations available on the project page, and observe that incorporating audio consistently improves performance during fine-tuning without compromising visual reasoning capabilities. These findings highlight that the audio signals in our dataset are not only well integrated, but also informative and complementary, establishing Valor32k-AVQA v2.0 as a valuable resource for developing and evaluating robust audio-visual question answering systems.

Zecheng Zhao, Selena Song, Tong Chen 0005, Zhi Chen 0010, Shazia Sadiq, Yadan Luo

Text-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synthetic videos perform in downstream tasks such as text-to-video retrieval (TVR). In this work, we introduce SynTVA, a new dataset and benchmark designed to evaluate the utility of synthetic videos for building retrieval models. Based on 800 diverse user queries derived from MSRVTT training split, we generate synthetic videos using state-of-the-art T2V models and annotate each video-text pair along four key semantic alignment dimensions: Object & Scene, Action, Attribute, and Prompt Fidelity. Our evaluation framework correlates general video quality assessment (VQA) metrics with these alignment scores, and examines their predictive power for downstream TVR performance. To explore pathways of scaling up, we further develop an Auto-Evaluator to estimate alignment quality from existing metrics. Beyond benchmarking, our results show that SynTVA is a valuable asset for dataset augmentation, enabling the selection of high-utility synthetic samples that measurably improve TVR outcomes. Project page and dataset can be found at https://jasoncodemaker.github.io/SynTVA/.

Peirong Zhang 0002, Yidan Zhang, Hanru Shi, Dianyu Wang, Xiaoxuan Liu, Lei Wang 0265

Referring multi-object tracking (RMOT), which aims to track one or more objects in a video based on a natural language query, is increasingly crucial for a wide range of real-world applications. However, the study of RMOT in satellite video (RMOT-SV) scenarios remains limited, largely due to the high cost of data acquisition and the difficulty of annotation. To address this gap, we introduce RefSat, the first dataset for benchmarking RMOT-SV. RefSat comprises 212 top-down viewpoint video clips, totaling 31,129 frames, collected from a variety of publicly available satellite video datasets. By combining manual annotations of object appearance and position with automatic motion estimation, we build a semi-automatic pipeline that generates high-quality natural language descriptions covering object attributes and motion trajectories, resulting in over 4,000 objects paired with carefully designed textual queries. RefSat features satellite-specific challenges such as small object sizes, cloud occlusions, and motion-referenced semantics. To address these, we introduce RSRefTrack, a tailored baseline designed for small object perception and motion-aware grounding, which outperforms existing state-of-the-art RMOT methods on the RefSat benchmark. Project page: https://github.com/Zhang-Peirong/RefSat

Yifan Wang, Jie Gui, Baosheng Yu, Qi Li 0005, Zhenan Sun, Juho Kannala, Guoying Zhao 0001

A major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we introduce FVeinSyn, a synthetic generator capable of producing diverse finger vein patterns with rich intra-class variations. Using FVeinSyn, we created FingerVeinSyn-5M -- the largest available finger vein dataset -- containing 5 million samples from 50,000 unique fingers, each with 100 variations including shift, rotation, scale, roll, varying exposure levels, skin scattering blur, optical blur, and motion blur. FingerVeinSyn-5M is also the first to offer fully annotated finger vein images, supporting deep learning applications in this field. Models pretrained on FingerVeinSyn-5M and fine-tuned with minimal real data achieve an average 53.91% performance gain across multiple benchmarks. The dataset is publicly available at: https://github.com/EvanWang98/FingerVeinSyn-5M.

Zhende Song, Chenchen Wang, Jiamu Sheng, Chi Zhang 0007, Shengji Tang, Jiayuan Fan 0001, Tao Chen 0003

Recent large vision-language models (LVLMs) for video understanding are primarily fine-tuned with various videos scraped from online platforms. Existing datasets, such as ActivityNet, require considerable human labor for structuring and annotation before effectively utilized for tuning LVLMs. While current LVLMs are primarily trained on existing datasets in broad, general-purpose settings, adapting them to specific downstream scenarios remains challenging, as collecting and annotating task-specific videos is highly labor-intensive and time-consuming. To address this issue, we propose a three-stage framework named DreamFrame for automatically generating style-consistent keyframes and corresponding question-answer (QA) pairs to support LVLM instruction tuning. DreamFrame generates datasets in a movie-like manner. First, we utilize an LLM to generate structured movie plots including movie prior information (like overview and style), frame descriptions and plot-related QA pairs, with a story expansion strategy to mitigate context length limitations. Then, to ensure visual consistency across generated frames, we design a Style Immobilization Process which maintains consistent style through an embedding learning strategy. Finally, frame descriptions and style embeddings are integrated to produce coherent keyframes. Using DreamFrame, we construct a dataset comprising approximately 1k stylized keyframe-like videos and 100k diverse QA pairs. Extensive fine-tuned experiments on various LVLM architectures demonstrate the effectiveness of the proposed dataset. Furthermore, based on the proposed dataset, we fine-tune a new LVLM named DreamFrame-7B, which significantly surpasses the previous similar-sized LVLMs (+2.2 compared with VideoLLaVA-7B on MvBench) across different benchmarks.

Zhixia Zhao, Qiyue Li 0001, Jie Li 0015, Richang Hong, Zhi Liu 0002

Gaussian splatting video has recently emerged as a promising representation for immersive 6-degree-of-freedom (6DoF) content due to its low-latency rendering, compact data structure, and high visual fidelity. In particular, 4D Gaussian splatting video-which models dynamic scenes as temporally evolving Gaussian splats in 3D space-offers an efficient solution for rendering photorealistic, interactive experiences. However, a systematic understanding of user behavior in such environments, especially head movement, remains largely unexplored due to the absence of dedicated datasets tailored to this format. This lack of data severely limits progress in viewpoint prediction, attention modeling, and video streaming optimization. To address this critical gap, we introduce ViewGauss-the first publicly available dataset that captures full 6DoF head movement during the viewing of 4D Gaussian splatting videos. Our dataset is collected from 35 participants using a high-precision Vive Focus Vision headset in a controlled environment, while they freely watched four reconstructed Gaussian splatting video sequences derived from the HiFi4G dataset. The data are recorded with high temporal resolution using position coordinates and unit quaternions, and organized into structured CSV files with precise timestamps for downstream synchronization and behavioral analysis. To demonstrate the practical value of ViewGauss, we conduct a preliminary viewpoint prediction experiment using the iTransformer model. The results show that head orientation patterns in 4D Gaussian splatting video scenes are not only temporally coherent but also learnable, highlighting the potential of ViewGauss as a benchmark for future behavioral modeling and predictive rendering systems. The dataset is publicly available at: https://github.com/Cedarleigh/ViewGauss-DataSet.

Wentao Mo, Qingchao Chen, Yuxin Peng 0001, Siyuan Huang 0001, Yang Liu 0105

The advancement of 3D vision-language (3D VL) learning is hindered by several limitations in existing 3D VL datasets: they rarely necessitate reasoning beyond a close range of objects in single viewpoint, and annotations often link instructions to single objects, missing richer contextual alignments between multiple objects. This significantly curtails the development of models capable of deep, multi-view 3D scene understanding over distant objects. To address these challenges, we introduce MV-ScanQA, a novel 3D question answering dataset where 68% of questions explicitly require integrating information from multiple views (compared to less than 7% in existing datasets), thereby rigorously testing multi-view compositional reasoning. To facilitate the training of models for such demanding scenarios, we present TripAlign dataset, a large-scale and low-cost 2D-3D-language pre-training corpus containing 1M <2D view, set of 3D objects, text> triplets that explicitly aligns groups of contextually related objects with text, providing richer, view-grounded multi-object multimodal alignment signals than previous single-object annotations. We further develop LEGO, a baseline method for the multi-view reasoning challenge in MV-ScanQA, transferring knowledge from pre-trained 2D LVLMs to 3D domain with TripAlign. Empirically, LEGO pre-trained on TripAlign achieves state-of-the-art performance not only on the proposed MV-ScanQA, but also on existing benchmarks for 3D dense captioning and question answering. Datasets and code are available at https://matthewdm0816.github.io/tripalign-mvscanqa.

Xiaoyu Guo, Pengzhi Zhong, Hao Zhang, Defeng Huang, Huikai Shao, Qijun Zhao, Shuiwang Li

Visual tracking has seen remarkable advancements, largely driven by the availability of large-scale training datasets that have enabled the development of highly accurate and robust algorithms. While significant progress has been made in tracking general objects, research on more challenging scenarios, such as tracking camouflaged objects, remains limited. Camouflaged objects, which blend seamlessly with their surroundings or other objects, present unique challenges for detection and tracking in complex environments. In critical fields like military, security, agriculture, and marine monitoring, accurately tracking camouflaged objects is essential. To address this gap, we introduce the Camouflaged Object Tracking Dataset (COTD), a specialized benchmark designed specifically for evaluating camouflaged object tracking methods. The COTD dataset comprises 200 sequences and approximately 80,000 frames, each annotated with detailed bounding boxes. Our evaluation of 20 existing tracking algorithms reveals significant deficiencies in their performance with camouflaged objects. To address these issues, we propose a novel tracking framework, HIPTrack-MLS, which demonstrates promising results in improving tracking performance for camouflaged objects. COTD and code are avialable at https://github.com/openat25/HIPTrack-MLS.

Mingsong Yang, Xinhong Hei 0001, Kehai Chen, Haining Meng, Haoyang Dong, Qin Zhao

Building Information Model (BIM) has become a significantly digital platform for representing buildings in the Architecture, Engineering, and Construction (AEC) industry. However, the absence of extensive, class- diverse, and balanced datasets at the BIM component level has limited the development of AI-driven BIM analysis. In this study, BIMCompNet is proposed as a large-scale multimodal dataset from Industry Foundation Classes (IFC), which can learn BIM component geometry features from multiple representation methods, including rendered views, point clouds, mesh structures, voxel grids, and semantic graphs. BIMCompNet is constructed by a standardized two-stage processing pipeline: (1) At the model level, geometry units are normalized to the SI units, models are converted to the IFC format, metadata is anonymized, and components are automatically extracted into individual IFC files. (2) At the component level, semantic labels are corrected, geometry and positioning are aligned, duplicates at model and project levels are removed, and five synchronized modalities (OBJ meshes, multi-view images, point clouds, voxel grids, and heterogeneous IFC graphs) are generated. BIMCompNet comprises 1,304,206 cleaned and labeled components across 87 IFC classes, collected from 1,607 real-world BIM models spanning 14 building types. To mitigate class imbalance, underrepresented classes are merged, and dominant classes are down-sampled to create balanced subsets suitable for robust AI model training and benchmarking. Benchmarking is performed on classification tasks by different models with multiple data modalities. Both the dataset and the processing pipeline will be publicly released to support reproducibility and private dataset extension.

Xing Zi, Jinghao Xiao, Yunxiao Shi, Xian Tao, Jun Li 0010, Ali Braytee, Mukesh Prasad

Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annotation richness, question diversity, and the assessment of specific reasoning capabilities. This paper introduces Remote Sensing Vision Language Model Question Answering (RSVLM-QA) dataset, a new large-scale, content-rich VQA dataset for the RS domain. RSVLM-QA is constructed by integrating data from several prominent RS segmentation and detection datasets: WHU, LoveDA, INRIA, and iSAID. We employ an innovative dual-track annotation generation pipeline. Firstly, we leverage Large Language Models (LLMs), specifically GPT-4.1, with meticulously designed prompts to automatically generate a suite of detailed annotations including image captions, spatial relations, and semantic tags, alongside complex caption-based VQA pairs. Secondly, to address the challenging task of object counting in RS imagery, we have developed a specialized automated process that extracts object counts directly from the original segmentation data; GPT-4.1 then formulates natural language answers from these counts, which are paired with preset question templates to create counting QA pairs. RSVLM-QA comprises 13,820 images and 162,373 VQA pairs, featuring extensive annotations and diverse question types. We provide a detailed statistical analysis of the dataset and a comparison with existing RS VQA benchmarks, highlighting the superior depth and breadth of RSVLM-QA's annotations. Furthermore, we conduct benchmark experiments on Six mainstream Vision Language Models (VLMs), demonstrating that RSVLM-QA effectively evaluates and challenges the understanding and reasoning abilities of current VLMs in the RS domain. We believe RSVLM-QA will serve as a pivotal resource for the RS VQA and VLM research communities, poised to catalyze advancements in the field. The dataset, generation code, and benchmark models are publicly available at https://github.com/StarZi0213/RSVLM-QA.

Liang Cheng, Hao Wang 0220, Chenwei Wu 0006, Haochen You, Xianhao Wu

Rainy weather typically leads to significantly reduced ambient illumination due to overcast skies.However, most existing image deraining datasets overlook this critical physical condition.They are usually constructed by linearly superimposing rain layers onto clean background images, without accounting for illumination degradation.This simplification introduces a clear domain gap between synthetic and real-world rainy images, thus limiting the generalization capability of current deraining algorithms.Moreover, existing methods predominantly focus on removing rain streaks while ignoring the simultaneous degradation caused by low-light conditions.To address these limitations, we introduce a new joint task: image deraining and low-light enhancement.Specifically, we construct a physically plausible dataset that simulates rainy scenes under low-light conditions, incorporating both rain streaks and raindrops with illumination-aware degradation modeling.In addition, we propose a baseline deraining network based on a multi-scale Mamba architecture, which jointly restores rain-free and well-lit images by effectively modeling both global illumination and local rain interference.Extensive experiments demonstrate that our method outperforms existing deraining approaches.The proposed dataset is released at https://drive.google.com/file/d/1QXxHqpYL7Q1TR5tdvAHm8tV2BdZgpGOc/view?usp=sharing.