OpenThoughts: Data Recipes for Reasoning Models
Stanford University, Anthropic · Harbor · Toyota Research Institute · University of California, Berkeley · University of Texas, Austin · University of California, Los Angeles · Juelich Supercomputing Center, LAION, Tuebingen University · Toyota Research Institute (TRI) · Google · New York University · UCLA · Stanford University · University of North Carolina at Chapel Hill · Department of Computer Science, University of Washington · Arizona State University · Princeton University · Georgia Institute of Technology · University of Washington · Computer Science Department, Stanford University · University of South Florida · Carnegie Mellon University · Anthropic · SynthLabs · McGill University · UNC Chapel Hill · University of Virginia Main Campus · Apple · Independent · Stanford University / NVIDIA · LAION; Juelich Supercomputing Center, Research Center Juelich · Technical University Munich · University of Southern California · Electrical Engineering & Computer Science Department, University of California, Berkeley · University of Washington / Stanford / Anthropic
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Reasoning models have made rapid progress on many benchmarks involving math, code, and science. Yet, there are still many open questions about the best train- ing recipes for reasoning since state-of-the-art models often rely on proprietary datasets with little to no public information available. To address this, the goal of the OpenThoughts project is to create open-source datasets for training reasoning models. Our OpenThoughts2-1M dataset led to OpenThinker2-32B, the first model trained on public reasoning data to match DeepSeek-R1-Distill-32B on standard reasoning benchmarks such as AIME and LiveCodeBench. We then improve our dataset further by systematically investigating each step of our data genera- tion pipeline with 1,000+ controlled experiments, which led to OpenThoughts3. Scaling the pipeline to 1.2M examples and using QwQ-32B as teacher yields our OpenThinker3-7B model, which achieves state-of-the-art results: 53% on AIME 2025, 51% on LiveCodeBench 06/24-01/25, and 54% on GPQA Dia- mond – improvements of 15.3, 17.2, and 20.5 percentage points compared to the DeepSeek-R1-Distill-Qwen-7B. All of our datasets and models are available on openthoughts.ai.