← 返回论文检索
NeurIPS 2024PosterAccept (Poster)

Fit for our purpose, not yours: Benchmark for a low-resource, Indigenous language

Suzanne Duncan, Gianna Leoni, Lee Steven, Keoni K Mahelona, Peter Lucas K Jones

Te Reo Irirangi o Te Hiku o Te Ika · Te Reo Irirangi o Te Hiku o te Ika

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Influential and popular benchmarks in AI are largely irrelevant to developing NLP tools for low-resource, Indigenous languages. With the primary goal of measuring the performance of general-purpose AI systems, these benchmarks fail to give due consideration and care to individual language communities, especially low-resource languages. The datasets contain numerous grammatical and orthographic errors, poor pronunciation, limited vocabulary, and the content lacks cultural relevance to the language community. To overcome the issues with these benchmarks, we have created a dataset for te reo Māori (the Indigenous language of Aotearoa/New Zealand) to pursue NLP tools that are ‘fit-for-our-purpose’. This paper demonstrates how low-resourced, Indigenous languages can develop tailored, high-quality benchmarks that; i. Consider the impact of colonisation on their language; ii. Reflect the diversity of speakers in the language community; iii. Support the aspirations for the tools they are developing and their language revitalisation efforts.