← 返回论文检索
NeurIPS 2024PosterAccept (Poster)

VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images

M. Maruf, Arka Daw, Kazi Sajeed Mehrab, Harish Babu Manogaran, Abhilash Neog, Medha Sawhney, Mridul Khurana, James Balhoff, Yasin Bakis, Bahadir Altintas, Matthew Thompson, Elizabeth Campolongo, Josef Uyeda, Hilmar Lapp, Henry Bart, Paula Mabee, Yu Su, Wei-Lun (Harry) Chao, Charles Stewart, Tanya Berger-Wolf, Wasila Dahdul, Anuj Karpatne

Virginia Tech · Oak Ridge National Laboratory (ORNL) · Virginia Polytechnic Institute & State University (Virginia Tech) · Virginia Polytechnic Institute and State University · Virginia Polytechnic Institute & State University (Virginia Tech) · University of North Carolina at Chapel Hill · Tulane University · Ohio State University, Columbus · The Ohio State University · Duke University · National Ecological Observatory Network, Battelle · Ohio State University (OSU) · Rensselaer Polytechnic Institute · Ohio State University · University of California, Irvine

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Images are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of $12$ state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of $469K$ question-answer pairs involving $30K$ images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images.