יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בחירת נתונים סינתטיים בהתאם לאימון

Training-Aware Target Coverage for Synthetic Data Selection
פותחים שיטה חדשה לבחירת נתונים סינתטיים לאימון מודלים. השיטה, TATC, בוחרת נתונים סינתטיים שישפרו את הביצועים של המודל במטלות ספציפיות. TATC נוסתה על מודל Qwen2.5-Math-1.5B-Instruct והראתה תוצאות טובות יותר משיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.00814v1 Announce Type: new Abstract: Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection
קרא במקור המקורי