כתבה
arXiv cs.LG ·
The Dynamics of Generalization in Deep Learning
תקציר מקורי באנגליתarXiv:2504.16450v4 Announce Type: replace Abstract: We derive a differential equation that governs the evolution of the generalization gap when a model is trained by gradient descent-based methods. This differential equation is driven by two key quantities, a contraction factor that brings together trajectories corresponding to slightly different datasets, and a perturbation factor that accounts for them training on different datasets. The coupled decay of contraction and perturbation guarantees a controlled accumulation of generalization gap during training. We analyze this differential equation to show that the generalization gap is given by a quadratic form that consists of an ``effective Gram matrix'' that depends upon the training trajectory and a certain residual of the predictor at
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית