כתבה
arXiv cs.LG ·
Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima
תקציר מקורי באנגליתarXiv:2607.08380v2 Announce Type: replace Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently violated in the training of deep neural networks. Recent work bridges this gap in the setting of overparametrised least-squares with a \emph{single scalar output}, providing a normal form for large-step GD in a neighbourhood of an \emph{isolated} flat minimum and establishing three corresponding convergence results. In this paper, we extend this theory in two directions: (1) to overparametrised least-squares with \emph{vector-valued outputs} (inc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית