כתבה
arXiv cs.LG ·
Mapping the Emergence of Regularization-Driven Dynamics in Grokking
תקציר מקורי באנגליתarXiv:2608.25813v3 Announce Type: replace Abstract: For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight decay (WD) perturbations across the pre-generalization plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering before visible generalization, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. Test-loss barriers between perturbed and base
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית