כתבה
arXiv cs.LG ·
A convergence result of a continuous model of deep learning via a \L{}ojasiewicz--Simon inequality
תקציר מקורי באנגליתarXiv:2311.15365v3 Announce Type: replace Abstract: We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space. The training dynamics are formulated as a Wasserstein-type gradient flow of an objective with a fixed $L^2$-regularization. Under suitable analyticity and growth assumptions, together with a coercivity assumption and sufficient regularity of the initial data, we prove that every curve of maximal slope converges to a single critical point of the objective as the training time tends to infinity. The proof combines compactness of the curve with a \L{}ojasiewicz--Simon inequality for the metric slope. To establish the inequality, we lift the object
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית