יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אאופיין נתונים סינתטיים באמצעות דינמיקת אימון

Synthetic Data Characterization via Training Dynamics
חוקרים אופיינו נתונים סינתטיים באמצעות דינמיקת אימון, ובדקו את השונות בין משפחות LLM וקנה מידה. הם גילו שאסטרטגיות בחירת נתונים משפיעות על מקורות נתונים שונים.
תקציר מקורי באנגליתarXiv:2609.39447v1 Announce Type: new Abstract: Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.
קרא במקור המקורי