כתבה
arXiv cs.LG ·
אסימפטוטיקה גבוה-ממדית ובחירת קבצים ללמידה עברית פרטית
High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning
במאמר זה, נוצרה שיטה לבחירת קבצים עבור למידה עברית פרטית, על ידי אסימפטוטיקה גבוה-ממדית. השיטה נותנת הבטחות פרטיות על תכונות סטטיסטיות ומאפשרת אופטימיזציה של פרמטרים.
תקציר מקורי באנגליתarXiv:2610.02578v1 Announce Type: cross Abstract: To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge est
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית