כתבה
arXiv cs.CL ·
הכל הדרכה: מתכון סינתטי יחיד ל-LLM
It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
חוקרים פיתחו את SYNTH, מאגר נתונים סינתטי פתוח ל-LLM. SYNTH מאפשר אימון מודלים ביעילות גבוהה יותר, עם פחות נתונים. המחקר בדק את היכולת של SYNTH לאמן מודלים כמו Monad ו-Baguettotron.
תקציר מקורי באנגליתarXiv:2609.37891v1 Announce Type: new Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית