MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
Yale University · Columbia University · Accenture · cleveland clinic foundation · Stevens Institute of Technology · Athens University of Economics and Business and Athena Research and Innovation Centre · University of Manchester · Athena Research and Innovation Center · National University of Singapore · University of Massachusetts Boston · University of Minnesota - Twin Cities · Augusta University · Yale University and Allen Institute for Artificial Intelligence · University of Manchester and The Fin AI · Harvard University · University of Florida · New York University · National Institute of Advanced Industrial Science and Technology · University of Montreal · Wuhan University
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.770 ↗
摘要
Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing evaluations of LLMs in finance remain text-only, monolingual, and largely saturated by current models. To bridge these gaps, we present MultiFinBen, the first expert-annotated multilingual (five languages) and multimodal (text, vision, audio) benchmark for evaluating LLMs in realistic financial contexts. MultiFinBen introduces two new task families: multilingual financial reasoning, which tests cross-lingual evidence integration from filings and news, and financial OCR, which extracts structured text from scanned documents containing tables and charts. Rather than aggregating all available datasets, we apply a structured, difficulty-aware selection based on advanced model performance, ensuring balanced challenge and removing redundant tasks. Evaluating 21 leading LLMs shows that even frontier multimodal models like GPT-4o achieve only 46.01% overall, stronger on vision and audio but dropping sharply in multilingual settings. These findings expose persistent limitations in multilingual, multimodal, and expert-level financial reasoning. All datasets, evaluation scripts, and leaderboards are publicly released.