כתבה
arXiv cs.LG ·
High-Probability Convergence of Clipped SGD under Heavy-Tailed Noise and $(L_0,L_1)$-Smoothness
תקציר מקורי באנגליתarXiv:2505.20817v3 Announce Type: replace-cross Abstract: Gradient clipping is widely used in language-model training to control heavy-tailed gradient noise and can improve convergence guarantees over stochastic gradient descent (SGD) under $(L_0,L_1)$-smoothness. Under these joint conditions, a central challenge is to obtain high-probability guarantees without exponential dependence on $L_1R_0$, where $R_0$ bounds the initial distance to a minimizer. We resolve this challenge for convex objectives, establishing, to the best of our knowledge, the first such guarantees for standard Clip-SGD. We assume unbiased stochastic gradients with bounded central $\alpha$-th moment, $\alpha\in(1,2]$. Our bounds have only polylogarithmic dependence on the inverse failure probability and recover known de
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית