כתבה
arXiv cs.LG ·
התכנסות שיטות גרדיאנט סטוכסטיות
Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H\"{o}lder Smoothness
חוקרים בדקו התכנסות שיטות גרדיאנט סטוכסטיות תחת רעש כבד-זנב וחלקות H"{o}lder. הם הראו כי שיטות אלו מתכנסות בקצב מהיר יותר משיטות קלאסיות. המחקר עלול לשפר את היכולת לאמן מודלים עמידים יותר.
תקציר מקורי באנגליתarXiv:2609.12785v1 Announce Type: new Abstract: Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice. In contrast, we study nonconvex stochastic optimization under the joint relaxation of these assumptions: objectives with $(L,s)$-H\"older continuous gradients, $s\in(0,1]$, and gradient noise satisfying only a bounded $\alpha$-th moment condition for $\alpha\in(1,2]$. We establish three convergence results. Firstly, that standard SGD converges at rate $O(T^{-s/(1+s)})$ whenever $\alpha\ge1+s$, extending the classical nonconvex SGD rate to heavy-tailed noise and H\"older smoothness simultaneously. Secondly, we analyze $\delta$-regularized gradient clipping ($\del
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית