יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

חשבון מידע מדויק לשיטות SGD

Exact information accounting for SGD methods
ניתוח מידע מדויק של שיטות SGD. הניתוח מראה כי צעד SGD מותנה הוא עדכון הממוצע האחורי של מודל בייס גאוסי. הניתוח חושף קשר בין שיפור וכישלון בהתאמה.
תקציר מקורי באנגליתarXiv:2610.00446v1 Announce Type: new Abstract: As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its variants. We show that a preconditioned SGD step is the posterior-mean update of a Gaussian Bayes model, and that its one-step regret splits into an intrinsic-time cost and a change in comparator information. The split extends to an identity for the objective itself. Convex convergence, strict-saddle-point escape, the link between flatness and generalization, the standard learning-rate schedules, adaptive optimizers, and the noisy, momentum, heavy-tailed, and gradient-free variants of SGD each correspond to a term or a special case of this identity. We measure its terms on synthetic and real train
קרא במקור המקורי