Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
University of Amsterdam · Google · Harvard University · New York University · Georgia Tech · University of Alberta · Sapienza University of Rome · Stanford University · Stanford · Universidad Complutense de Madrid · Carnegie Mellon University · xAI · Koç University · Google DeepMind · Allen Institute for AI · University of California Berkeley · Allen Institute for Artificial Intelligence · Google, Inc. · Synapc · Google Research · Technion · University of Illinois, Urbana-Champaign · Columbia University · Tencent AI Lab Seattle · Princeton University · University of Oxford · Google Deep Mind · Google Brain · INRIA · Tel Aviv University · Microsoft Research · Winterlight Labs · University of Virginia · Arizona State University · Heidelberg Institute for Theoretical Studies · Salesforce AI Research · Cornell University · UNC Chapel Hill and Hugging Face · OpenAI · University of Maryland, College Park · Brown University · UCI · University of California, Los Angeles · Rutgers University · University of Kassel, hessian.AI, and ScaDS.AI · University of Notre Dame · 3M Health Information Systems Inc. · Vrije Universiteit Amsterdam · University of California, San Diego · University of California, Irvine · Georgia Institute of Technology · FAIR · Elicit · Thapar University · Amazon · National University of Science and Technology · None · Yale University · Bauhaus Universität Weimar · Apple · Queen's University · Universitat Politècnica de València · Forschungszentrum Jülich · SAE Expression College · Research, Google · University of Michigan - Ann Arbor · Anthropic · University of Edinburgh · Örebro University · Koc University · Soleda AI · Department of Computer Science, University of Toronto · Google AI · Umea University · University of California, Berkeley · Meta AI · Hudson River Trading · Strathmore University · Technion - Israel Institute of Technology, Technion · Universidad Politécnica de Valencia · UIUC · Valencian Research Institute for Artificial Intelligence · Karlsruher Institut für Technologie · University of Cambridge · Max-Planck Institute for Human Development · Swiss Federal Institute of Technology · Computer Science Department, Stanford University · University of Groningen · Wroclaw University of Science and Technology · Facebook · Universitat Politecnica de Valencia · California State University, Chico · Drexel University · University of Pennsylvania · Heidelberg University · Emory University · EleutherAI · Stevens Institute of Technology · The Center for Information and Language Processing · UNIVERSITAT POLITÈCNICA DE VALÈNCIA · UNC Chapel Hill · University of Bristol · Department of Computer Science, ETHZ - ETH Zurich · Colorado School of Mines · Friedrich-Schiller Universität Jena · USC/ISI · Samaya AI · NVIDIA · University of Kassel · University of Southern California · Cruise · State University of New York at Stony Brook · General Agents · AWS AI Labs · University of Memphis · Meta · School of Computer Science, Carnegie Mellon University · Google Deepmind · MIT · National Research Council Canada · Indian Institute of Technology Madras · Rice University · Eberhard-Karls-Universität Tübingen · Bloomberg · School of Engineering and Applied Science, University of Pennsylvania · Siemens Healthineers · Hasso Plattner Institute · University of Pennsylvania, University of Pennsylvania · Boston University, Boston University · University of Washington/AI2 · Lone Star College System · UC Berkeley · FAR AI · Cisco · Universität Potsdam · ex-Allen Institute for Artificial Intelligence (now DeepMind) · Microsoft · METR · Institute for Logic, Language and Computation, University of Amsterdam · Ecole Normale Supérieure de Paris · IT:U Interdisciplinary Transformation University Austria · University of Chicago · Capital One · TTI-Chicago · Alignment Research Center · Microsoft Research Montreal · McGill University / Mila · University of North Carolina -Chapel Hill · Bytedance · Pennsylvania State University · ML Collective · Naver Labs Europe · Amirkabir University of Technology · Hahn Air GmbH · University of Toronto · University of Massachusetts Amherst · CMU, Carnegie Mellon University · Sapienza & IST Austria · NYU · SalesForce.com · National University of Singapore · Charles River · Alphabet · Center for Information and Language Processing · The Hong Kong University of Science and Technology · Oracle and the University of Pennsylvania · Google Research (Brain) · University College London · CMU · University of Wisconsin, Madison · University of Montreal · University of Utah · Toyota Technological Institute at Chicago · Max-Planck-Institute for Intelligent Systems, Max-Planck Institute · University of Washington · University of Hong Kong · ML Collective; Google DeepMind · Max Planck Institute for Intelligent Systems and ETH Zurich · Hacettepe University · Saarland University · University of Michigan Ann Arbor · Johns Hopkins University · Apple AI/ML
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyond the Imitation Game benchmark (BIG- bench). BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions. Task topics are diverse, drawing problems from linguistics, childhood develop- ment, math, common-sense reasoning, biology, physics, social bias, software development, and beyond. BIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models. We evaluate the behavior of OpenAI's GPT models, Google- internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters. In addition, a team of human expert raters performed all tasks in order to provide a strong baseline. Findings include: model performance and calibration both improve with scale, but are poor in absolute terms (and when compared with rater performance); performance is remarkably similar across model classes, though with benefits from sparsity; tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit "breakthrough" behavior at a critical scale often involve multiple steps or components, or brittle metrics; social bias typically increases with scale in settings with ambiguous context, but this can be improved with prompting.