OmniBench: Towards The Future of Universal Omni-Language Models
The University of Manchester · University of Michigan - Ann Arbor · Centre for Digital Music, Queen Mary University of London · Carnegie Mellon University · Guangdong OPPO Mobile Telecommunications Corp.,Ltd. · Alibaba Group · University of the Chinese Academy of Sciences · Nanjing University · Nanjing University of Science and Technology · University of Manchester · Queen Mary, University of London · National University of Singapore · China University of Geoscience Beijing · Northwest Polytechnical University Xi'an · nanjing university · Chinese Academy of Sciences, China · Google DeepMind · Queen Mary University of London · Key Laboratory of Machine Perception
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Recent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as omni-language models (OLMs). OmniBench features high-quality human annotations that require integrated understanding across all modalities. Our evaluation reveals that: i) open-source OLMs show significant limitations in instruction-following and reasoning in tri-modal contexts; and ii) most baseline models perform poorly (below 50% accuracy) even with textual alternatives to image/audio inputs. To address these limitations, we develop OmniInstruct, an 96K-sample instruction tuning dataset for training OLMs. We advocate for developing more robust tri-modal integration techniques and training strategies to enhance OLM performance. Codes and data could be found at https://m-a-p.ai/OmniBench/.