← 返回论文检索
EMNLP 2024emnlpfindings

“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions

Mihai Masala, Denis Ilie-Ablachim, Alexandru Dima, Dragos Georgian Corlatescu, Miruna-Andreea Zavelca, Ovio Olaru, Simina-Maria Terian, Andrei Terian, Marius Leordeanu, Horia Velicu, Marius Popescu, Mihai Dascalu, Traian Rebedea

The Institute of Mathematics of the Romanian Academy (IMAR) · The National University of Science and Technology Politehnica Bucharest · CrowdStrike and University Politehnica of Bucharest · University of Bucharest · University Lucian Blaga of Sibiu · Lucian Blaga University of Sibiu · Norwegian Research Center (NORCE), University Politehnica of Bucharest and Institute of Mathematics of the Romanian Academy · University Politehnica of Bucharest · NVIDIA and University Politehnica of Bucharest

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2024.findings-emnlp.681 ↗

摘要

In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English greatly exceeds other languages. To our knowledge, we are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate, and release open-source LLMs tailored for Romanian. We evaluate our methods on four different categories, including academic benchmarks, MT-Bench (manually translated), and a professionally built historical, cultural, and social benchmark adapted to Romanian. We argue for the usefulness and high performance of RoLLMs by obtaining state-of-the-art results across the board. We publicly release all resources (i.e., data, training and evaluation code, models) with the goal of supporting and encouraging research on Romanian LLMs while concurrently creating a generalizable recipe adequate for other low or less-resourced languages.