כתבה
arXiv cs.LG ·
הבטחות כלליות לכלליות: טונינג נתונים-נתונים של גרדיאנט דסנט עם עדכוני לאנג'בין
Generalization Guarantees on Data-Driven Tuning of Gradient Descent with Langevin Updates
במאמר זה, המחברים חקרו את הלמידה ללמוד דרך עדכון גרדיאנט דסנט. הם הציגו אלגוריתם של גרדיאנט דסנט עם עדכוני לאנג'בין, שמעריך את הממוצע של ההפצה האחרונה של ההפסד. הם הוכיחו את קיום ההגדרה האופטימלית של הפרמטרים האופטימליים, והציגו הבטחות כלליות לכלליות במסגרת הטונינג הנתונים-נתונים.
תקציר מקורי באנגליתarXiv:2604.13130v2 Announce Type: replace Abstract: We study learning to learn through the lens of hyperparameter tuning. We propose the Langevin Gradient Descent Algorithm (LGD), which approximates the mean of the posterior distribution defined by the loss function and regularizer of a regression task with convex objective. For classification tasks, the LGD algorithm estimates the posterior probabilities of each class on the test set. We prove the existence of an optimal hyperparameter configuration for which the LGD algorithm achieves the Bayes' optimal solution for squared loss on regression tasks, and for which LGD closely approximates the posterior probabilities for well-specified classification tasks. Subsequently, we study generalization guarantees on meta learning optimal hyperpara
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית