כתבה
arXiv cs.LG ·
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
תקציר מקורי באנגליתarXiv:2605.17849v2 Announce Type: replace-cross Abstract: LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound regime does not mean the model has fully utilized its organic corpus. In this paper, we introduce SynPro, a synthetic data generation framework that helps LLMs more thoroughly learn from limited organic data. SynPro applies two operations, rephrasing and reformatting, that present the same organic source in diverse forms to facilitate deeper learning without introducing external information. Both generators are optimized via reinforcement learning with quality, faithfulness, and data influence rewards, and are continuously updated as pretraining plateaus to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית