כתבה
arXiv cs.AI ·
למידה עם נתונים סינתטיים באמצעות SGD בתצורת רגרסיה לינארית גבוהה
Learning with Synthetic Data via SGD in High-Dimensional Linear Regression
נתונים סינתטיים גורמים להשפעה על הכלליות של המודל ברגרסיה לינארית גבוהה. המחקר חוקר את השפעת הנתונים הסינתטיים על הכלליות של המודל, ומציע פתרונות לשיפור הכלליות של המודל.
תקציר מקורי באנגליתarXiv:2609.09572v1 Announce Type: cross Abstract: Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Dohmatob et al., 2024), where any fixed fraction of synthetic data prevents model performance from improving under data scaling, leaving a non-vanishing excess risk floor. In this paper, we study how synthetic data affects the generalization of one-pass SGD in high-dimensional linear regression with model shift. We establish finite-sample risk bounds for mixed and two-stage training, separating standard bias and variance from source-mismatch effects, namely fluctuation and persistent drift under mixing and filtered initialization bias under two-stage. These bounds reveal a sharp contrast: mixed
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית