יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

RePro: טירונות מודלי לשפה להחזרת האינטרנט באופן אמין לטירונות

RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
RePro: טירונות מודלי לשפה להחזרת האינטרנט באופן אמין לטירונות. המחברים הציגו טכניקה חדשה לטירונות מודלי לשפה שמטרתה להחזרת נתונים מהאינטרנט באופן אמין. הם טענו כי טכניקה זו יכולה לשפר את יעילות הטירונות ולהפחית את עלויות הטירונות.
תקציר מקורי באנגליתarXiv:2510.10681v2 Announce Type: replace-cross Abstract: High-quality data is a cornerstone of large language model (LLM) pretraining, yet its growth has not kept pace with the needs of frontier models. In this paper, we introduce RePro, a novel web recycling method that trains a relatively small LM with reinforcement learning to generate effective and faithful rephrasings of pretraining data. Specifically, we design one quality reward and three faithfulness rewards, optimizing the LM rephraser to convert organic data into high-quality rephrasings while maintaining its core semantics and structure. In our experiment, we train a rephraser as small as 1B parameters to recycle 72B tokens sampled from DCLM-RefinedWeb. Pretraining results on 400M, 1.4B, and 2.8B models demonstrate that RePro d
קרא במקור המקורי