← 返回论文检索
NeurIPS 2023Spotlight PosterAccept (Spotlight)

Renku: a platform for sustainable data science

Rok Roškar, Chandrasekhar Ramakrishnan, Michele Volpi, Fernando Perez-Cruz, Lilian Gasser, Firat Ozdemir, Patrick Paitz, Mohammad Alisafaee, Philipp Fischer, Ralf Grubenmann, Eliza Harris, Tasko Olevski, Carl Remlinger, Luis Salamanca, Elisabet Capon Garcia, Lorenzo Cavazzi, Jakub Chrobasik, Darlin Cordoba Osnas, Alessandro Degano, Jimena Dupre, Wesley Johnson, Eike Kettner, Laura Kinkead, Sean Murphy, Flora Thiebaut, Olivier Verscheure

Swiss Data Science Center / ETH Zürich · ETHZ - ETH Zurich · ETH Zurich · Swiss Data Science Center, ETH Zurich and EPFL · Swiss Federal Institute for Forest, Snow and Landscape Research WSL · University of Tehran, University of Tehran · ETHZ - ETH Zürich · FHNW - Fachhochschule Nordwestschweiz · University of Toronto · Université Gustave Eiffel · Swiss Data Science Center, ETH Zürich · Swiss Data Science Center · Silesian Technical University of Gliwice · Universidad del Valle del Cauca · University of Turin · ZHdK - Zurich University of the Arts · Friedrich-Schiller Universität Jena · École Polytechnique · Swiss Data Science Center EPFL & ETH Zurich

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Data and code working together is fundamental to machine learning (ML), but the context around datasets and interactions between datasets and code are in general captured only rudimentarily. Context such as how the dataset was prepared and created, what source data were used, what code was used in processing, how the dataset evolved, and where it has been used and reused can provide much insight, but this information is often poorly documented. That is unfortunate since it makes datasets into black-boxes with potentially hidden characteristics that have downstream consequences. We argue that making dataset preparation more accessible and dataset usage easier to record and document would have significant benefits for the ML community: it would allow for greater diversity in datasets by inviting modification to published sources, simplify use of alternative datasets and, in doing so, make results more transparent and robust, while allowing for all contributions to be adequately credited. We present a platform, Renku, designed to support and encourage such sustainable development and use of data, datasets, and code, and we demonstrate its benefits through a few illustrative projects which span the spectrum from dataset creation to dataset consumption and showcasing.