CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Common Crawl Foundation · EleutherAI · Factored · University of Pretoria · Lelapa AI and University of Pretoria · Imperial College London and Bayero University, Kano-Nigeria · Department of Computer Science, University College London · King Saud University · University of Michigan - Ann Arbor and USIU- Africa · Universiti Teknologi Petronas · EPFL · Khatam University, Tehran Institute for Advanced Studies (TEIAS) · Universitas Indonesia · Universität Trier and Turing · Michigan State University · National Taiwan University of Science and Technology and Academia Sinica · University of Bath · University of Copenhagen · Independent Researcher · University of Edinburgh, University of Edinburgh · Massachusetts Institute of Technology and International Business Machines · Stanford University · Prince Sattam bin Abdulaziz University and Benha University · Universität Hamburg · SEACrowd and Electric Sheep · Computational Intelligence and Operations Laboratory · IT University of Copenhagen · Ofis Publik ar Brezhoneg · Oracle · Johns Hopkins University · University of New Haven · Technische Universität München · Stevens Institute of Technology · Cerebras Systems, Inc · Capital One · Carnegie Mellon University · Universidad de Zaragoza · German Research Center for AI · University of Ibadan · Cariva · National Research Council Canada · Lam Research · NEC · NeuralShift · independent · Hasso Plattner Institute · Independent · Mohamed bin Zayed University of Artificial Intelligence · niversity of Technology Nuremberg and Universität Mannheim · Thammasat University · INRIA and German Research Center for AI · INRIA · Nasarawa State University Keffi · Meta Fundamental Artificial Intelligence Research (FAIR) · Inria
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.1527 ↗
摘要
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper, we introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. Many of the included languages have been previously under-served, making CommonLID a key resource for developing more representative high-quality text corpora. We show CommonLID’s value by using it, alongside five other common evaluation sets, to test eight popular LID models. We analyse our results to situate our contribution and to provide an overview of the state of the art. In particular, we highlight that existing evaluations overestimate LID accuracy for many languages in the web domain. We make CommonLID and the code used to create it available under an open, permissive license.